SearcharxivSearch

arXiv subjects

Boyang Shen

Publications and source records attributed to Boyang Shen.

8 recordsLinked to original sources

LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models

Current Vision-Language-Action (VLA) models typically treat the deepest representation of a vision-language backbone as universally optimal for action prediction. However, robotic manipulation is composed of many frequent closed-loop spatial adjustments, for which excessive abstraction may waste computation and weaken low-level geometric cues essential for precise control. Existing early-exit strategies attempt to reduce computation by stopping at predefined layers or applying heuristic rules such as action consistency, but they do not directly answer when a representation is actually sufficient for action. In this paper, we present LoopVLA, a recurrent VLA architecture that jointly learns representation refinement, action prediction, and sufficiency estimation. LoopVLA iteratively applies a shared Transformer block to refine multimodal tokens, and at each iteration produces both a candidate action and a sufficiency score that estimates whether further refinement is necessary. By sharing parameters across iterations, LoopVLA decouples refinement from absolute layer indices and grounds sufficiency estimation in the evolving representation itself. Since sufficiency has no direct supervision, we introduce a self-supervised distribution alignment objective, where intermediate confidence scores are trained to match the relative action quality across refinement steps, thereby linking sufficiency learning to policy optimization signals. Experiments on LIBERO, LIBERO-Plus, and VLA-Arena show that LoopVLA pushes the efficiency-performance frontier of VLA policies, reducing parameters by 45% and improving inference throughput by up to 1.7 times while matching or outperforming strong baselines in task success.

cs.AI

FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image Editing

Text-guided image editing has advanced rapidly with the rise of diffusion models. While flow-based inversion-free methods offer high efficiency by avoiding latent inversion, they often fail to effectively integrate source information, leading to poor background preservation, spatial inconsistencies, and over-editing due to the lack of effective integration of source information. In this paper, we present FIA-Edit, a novel inversion-free framework that achieves high-fidelity and semantically precise edits through a Frequency-Interactive Attention. Specifically, we design two key components: (1) a Frequency Representation Interaction (FRI) module that enhances cross-domain alignment by exchanging frequency components between source and target features within self-attention, and (2) a Feature Injection (FIJ) module that explicitly incorporates source-side queries, keys, values, and text embeddings into the target branch's cross-attention to preserve structure and semantics. Comprehensive and extensive experiments demonstrate that FIA-Edit supports high-fidelity editing at low computational cost (~6s per 512 * 512 image on an RTX 4090) and consistently outperforms existing methods across diverse tasks in visual quality, background fidelity, and controllability. Furthermore, we are the first to extend text-guided image editing to clinical applications. By synthesizing anatomically coherent hemorrhage variations in surgical images, FIA-Edit opens new opportunities for medical data augmentation and delivers significant gains in downstream bleeding classification. Our project is available at: https://github.com/kk42yy/FIA-Edit.

cs.CV

The SAGES Critical View of Safety Challenge: A Global Benchmark for AI-Assisted Surgical Quality Assessment

Advances in artificial intelligence (AI) for surgical quality assessment promise to democratize access to expertise, with applications in training, guidance, and accreditation. This study presents the SAGES Critical View of Safety (CVS) Challenge, the first AI competition organized by a surgical society, using the CVS in laparoscopic cholecystectomy, a universally recommended yet inconsistently performed safety step, as an exemplar of surgical quality assessment. A global collaboration across 54 institutions in 24 countries engaged hundreds of clinicians and engineers to curate 1,000 videos annotated by 20 surgical experts according to a consensus-validated protocol. The challenge addressed key barriers to real-world deployment in surgery, including achieving high performance, capturing uncertainty in subjective assessment, and ensuring robustness to clinical variability. To enable this scale of effort, we developed EndoGlacier, a framework for managing large, heterogeneous surgical video and multi-annotator workflows. Thirteen international teams participated, achieving up to a 17% relative gain in assessment performance, over 80% reduction in calibration error, and a 17% relative improvement in robustness over the state-of-the-art. Analysis of results highlighted methodological trends linked to model performance, providing guidance for future research toward robust, clinically deployable AI for surgical quality assessment.

cs.CV

Learning-Based Surrogate Method for Stochastic Optimization under Decision-Dependent Uncertainty with Adaptive Random Designs

We study stochastic programs in which the latent decision-dependent uncertainty is described via a nonparametric regression model. The major challenge is that, without convexity assumptions on either the cost function or the regression model, the resulting objective is both nonconvex and nonsmooth, and its first-order information is unavailable due to the unknown decision-dependent distribution. To address this issue, we construct a learning-based surrogate model that integrates simulation and statistical learning by embedding Jacobian estimates of the regression function, which are updated iteratively and interactively during the optimization procedure. We develop an adaptive random design that concentrates design points around the current iterate for Jacobian estimation and we show that the mean squared error of Jacobian estimates achieves a dimension-independent convergence rate. Building on this, we propose the learning-based stochastic prox-linear (L-SPL) algorithm with adaptive random design and establish its nonasymptotic convergence rates under various parameter settings. Numerical results demonstrate that L-SPL algorithm significantly improves sample efficiency and achieves substantially lower objective values compared to the state-of-the-art methods. More broadly, our method implies that the statistical design in an iterative learning-based optimization algorithm can be novelly tailored to the local information of the optimization procedure to sharpen estimates and enhance the convergence performance and sample efficiency of the resulting algorithm.

math.OC

JobViz: Skill-driven Visual Exploration of Job Advertisements

Online job advertisements on various job portals or websites have become the most popular way for people to find potential career opportunities nowadays. However, the majority of these job sites are limited to offering fundamental filters such as job titles, keywords, and compensation ranges. This often poses a challenge for job seekers in efficiently identifying relevant job advertisements that align with their unique skill sets amidst a vast sea of listings. Thus, we propose well-coordinated visualizations to provide job seekers with three levels of details of job information: a skill-job overview visualizes skill sets, employment posts as well as relationships between them with a hierarchical visualization design; a post exploration view leverages an augmented radar-chart glyph to represent job posts and further facilitates users' swift comprehension of the pertinent skills necessitated by respective positions; a post detail view lists the specifics of selected job posts for profound analysis and comparison. By using a real-world recruitment advertisement dataset collected from 51Job, one of the largest job websites in China, we conducted two case studies and user interviews to evaluate JobViz. The results demonstrated the usefulness and effectiveness of our approach.

cs.HC

Review of the AC Loss Computation for HTS using the H-formulation

This article presents a review of the finite element method (FEM) model based on the $H$ formulation of Maxwell's equations used to calculate AC losses in high temperature superconductor (HTS) tapes, cables and windings for different applications. This model, which uses the components of the magnetic field as state variables, has been gaining a great popularity and has been in use in tens of research groups around the world. This contribution first reviews the equations on which the model is based and their implementation in finite element method programs for different cases, such 2D longitudinal and axis-symmetric geometries, 3D geometries. Modeling strategies to tackle large number of HTS tapes, such as multi-scale and homogenization methods, are also introduced. Then, the second part of the article reviews the applications for which the $H$ formulations has been used to calculate AC losses, ranging from individual tapes, to complex cables and large magnet windings. Afterwards, a section is dedicated to the discussion of the $H$ formulation in terms of computational efficiency and easiness of implementation. Its pros and cons are listed. Finally, the last section draws the main conclusions.

cond-mat.supr-con

A kilo-Ampere level HTS flux pump

This paper reports a newly developed high current transformer-rectifier High-Tc Superconducting (HTS) flux pump switched by dynamic resistance. A quasi-persistent current of over 1.1 kA has been achieved at 77 K using the device, which is the highest reported operating current by any HTS flux pumps to date. The size of the device is much smaller than traditional current leads and power supplies at the same current level. Parallel YBCO coated conductors are used in the transformer secondary winding as well as in the superconducting load coil to achieve high current. The output current is limited by the critical current of the load rather than the flux pump itself. Moreover, at over 1 kA current level, the device can maintain high flux injection accuracy, and the overall flux ripple is less than 0.2 mili-Weber. The work has shown the potential of using the device to operate high field HTS magnets in ultra-high quasi-persistent current mode, thus substantially reducing the inductance, size, weight, and cost of high field magnets, making them more accessible. It also indicates that the device is promising for powering HTS NMR/MRI magnets, in which the requirement for magnetic field satiability is demanding.

physics.app-ph

Investigation of AC Loss in HTS Cross-Conductor Cables for Electrical Power Transmission

This paper presents the alternating current (AC) loss analysis on high-temperature superconductor (HTS) Cross-Conductor (CroCo) cables, in order to evaluate whether they could be utilized for electrical power transmission. The modeling of HTS CroCo cables was based on a cable assembled at the Karlsruhe Institute of Technology (KIT) and the AC loss calculation was based on the H-formulation model implemented in the finite-element method (FEM) software package COMSOL Multiphysics. The AC loss calculations have been carried out for isolated single-phase CroCo cable and three-phase CroCo cables. The AC loss angular dependence of a particular phase of CroCo cables during three phase operation has been studied. The current distributions of individual tapes within CroCo cables have been investigated.

cond-mat.supr-con