SearcharxivSearch

arXiv subjects

Xinmin Liu

Publications and source records attributed to Xinmin Liu.

12 recordsLinked to original sources

LIBERO-X: Robustness Litmus for Vision-Language-Action Models

Reliable benchmarking is critical for advancing Vision-Language-Action (VLA) models, as it reveals their generalization, robustness, and alignment of perception with language-driven manipulation tasks. However, existing benchmarks often provide limited or misleading assessments due to insufficient evaluation protocols that inadequately capture real-world distribution shifts. This work systematically rethinks VLA benchmarking from both evaluation and data perspectives, introducing LIBERO-X, a benchmark featuring: 1) A hierarchical evaluation protocol with progressive difficulty levels targeting three core capabilities: spatial generalization, object recognition, and task instruction understanding. This design enables fine-grained analysis of performance degradation under increasing environmental and task complexity; 2) A high-diversity training dataset collected via human teleoperation, where each scene supports multiple fine-grained manipulation objectives to bridge the train-evaluation distribution gap. Experiments with representative VLA models reveal significant performance drops under cumulative perturbations, exposing persistent limitations in scene comprehension and instruction grounding. By integrating hierarchical evaluation with diverse training data, LIBERO-X offers a more reliable foundation for assessing and advancing VLA development.

cs.CV

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, this report presents a comprehensive assessment of their frontier risks. Drawing on the E-T-C analysis (deployment environment, threat source, enabling capability) from the Frontier AI Risk Management Framework (v1.0) (SafeWork-F1-Framework), we identify critical risks in seven areas: cyber offense, biological and chemical risks, persuasion and manipulation, uncontrolled autonomous AI R\&D, strategic deception and scheming, self-replication, and collusion. Guided by the "AI-$45^\circ$ Law," we evaluate these risks using "red lines" (intolerable thresholds) and "yellow lines" (early warning indicators) to define risk zones: green (manageable risk for routine deployment and continuous monitoring), yellow (requiring strengthened mitigations and controlled deployment), and red (necessitating suspension of development and/or deployment). Experimental results show that all recent frontier AI models reside in green and yellow zones, without crossing red lines. Specifically, no evaluated models cross the yellow line for cyber offense or uncontrolled AI R\&D risks. For self-replication, and strategic deception and scheming, most models remain in the green zone, except for certain reasoning models in the yellow zone. In persuasion and manipulation, most models are in the yellow zone due to their effective influence on humans. For biological and chemical risks, we are unable to rule out the possibility of most models residing in the yellow zone, although detailed threat modeling and in-depth assessment are required to make further claims. This work reflects our current understanding of AI frontier risks and urges collective action to mitigate these challenges.

cs.AI

Programmable Cycle-Specified Queue for Long-Distance Industrial Deterministic Packet Scheduling

The time-critical industrial applications pose intense demands for enabling long-distance deterministic networks. However, previous priority-based and weight-based scheduling methods focus on probabilistically reducing average delay, which ignores strictly guaranteeing task-oriented on-time packet delivery with bounded worst-case delay and jitter. This paper proposes a new Programmable Cycle-Specified Queue (PCSQ) for long-distance industrial deterministic packet scheduling. By implementing the first high-precision rotation dequeuing, PCSQ enables microsecond-level time slot resource reservation (noted as T) and especially jitter control of up to 2T. Then, we propose the cycle tags computation to approximate cyclic scheduling algorithms, which allows packets to actively pick and lock their favorite queue in a sequence of nodes. Accordingly, PCSQ can precisely defer packets to any desired time. Further, the queue coordination and cycle mapping mechanisms are delicately designed to solve the cycle-queue mismatch problem. Evaluation results show that PCSQ can schedule tens of thousands of time-sensitive flows and strictly guarantee $ms$-level delay and us-level jitter.

cs.NI

Affordances-Oriented Planning using Foundation Models for Continuous Vision-Language Navigation

LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) task. However, existing LLM-based methods often focus only on solving high-level task planning by selecting nodes in predefined navigation graphs for movements, overlooking low-level control in navigation scenarios. To bridge this gap, we propose AO-Planner, a novel Affordances-Oriented Planner for continuous VLN task. Our AO-Planner integrates various foundation models to achieve affordances-oriented low-level motion planning and high-level decision-making, both performed in a zero-shot setting. Specifically, we employ a Visual Affordances Prompting (VAP) approach, where the visible ground is segmented by SAM to provide navigational affordances, based on which the LLM selects potential candidate waypoints and plans low-level paths towards selected waypoints. We further propose a high-level PathAgent which marks planned paths into the image input and reasons the most probable path by comprehending all environmental information. Finally, we convert the selected path into 3D coordinates using camera intrinsic parameters and depth information, avoiding challenging 3D predictions for LLMs. Experiments on the challenging R2R-CE and RxR-CE datasets show that AO-Planner achieves state-of-the-art zero-shot performance (8.8% improvement on SPL). Our method can also serve as a data annotator to obtain pseudo-labels, distilling its waypoint prediction ability into a learning-based predictor. This new predictor does not require any waypoint data from the simulator and achieves 47% SR competing with supervised methods. We establish an effective connection between LLM and 3D world, presenting novel prospects for employing foundation models in low-level motion control.

cs.RO

Revisiting the measurements and interpretations of DLVO forces

The DLVO theory and electrical double layer (EDL) theory are the foundation of colloid and interface science. With the invention and development of surface forces apparatus (SFA) and atomic force microscope (AFM), the measurements and interpretations of DLVO forces (i.e., mainly measuring the EDL force (electrostatic force) FEDL and van der Waals force FvdW, and interpreting the potential ψ, charge density σ, and Hamaker constant H) can be greatly facilitated by various surface force measurement techniques, and would have been very promising in advancing the DLVO theory, EDL theory, and colloid and interface science. However, although numerous studies have been conducted, pervasive anomalous results can be identified throughout the literature, main including: (1) the fitted ψ/σ is normally extremely small (ψ can be close to or (much) smaller than ψζ (zeta potential)) and varies greatly; (2) the fitted ψ/σ can exceed the allowable range of calculation; and (3) the measured FvdW and the fitted H vary greatly. Based on rigorous and comprehensive arguments, we have reasonably explained the pervasive anomalous results in the literature and further speculated that, the pervasive anomalous results are existing but not noticed and questioned owing to the two important aspects: (1) the pervasive unreasonable understandings of EDL theory and (2) the commonly neglected systematic errors. Consequently, we believe that the related studies have been seriously hampered. We therefore call for re-examination and re-analysis of related experimental results and theoretical understandings by careful consideration of the EDL theory and systematic errors. On these bases, we can interpret the experimental results properly and promote the development of EDL theory, colloid and interface science, and many related fields.

physics.chem-ph

Full-body deep learning-based automated contouring of contrast-enhanced murine organs for small animal irradiator CBCT

Purpose. To alleviate the manual contouring burden, deep learning (DL) based automated contouring has been explored. However, due to the poor contrast resolution of preclinical irradiator CBCT, these methods have been limited to high contrast - minimally anatomically complex - structures such as the heart and lungs. Thus, low contrast abdominal CBCT DL-based segmentation has yet to be addressed. In this work we explore a DL-based model in conjunction with iodine-based contrast agent approach to allow precise automatic contouring of mouse abdominal, thorax, and skeletal structures in under a second. Methods. A DL U-net-like architecture was trained to contour mice organs in small animal radiation research platform CBCT scans. 41 mice were contoured by a human expert, using semi-automatic segmentation methods, after injection of iodine contrast agent, establishing a ground truth for the DL model. The model was trained on a dataset of 26 mice, while 2 mice were used for validation, tuning the model during training, and 15 mice used for performance evaluation testing. The model consists of a pre-processor, and a post-processor for volumetric reconstruction of the DL-predicted probability maps. Model performance was evaluated using both qualitative and distance metrics, including the dice similarity score, precision score, Hausdorff Distance (HD), and mean surface distance (MSD). Results. Performance of the DL-based iodine contrast-enhanced model provided high quality predicted contours in under a second, with the median for all organs being reported: dice $>$ 91\%, precision $>$ 95\%, HD50 $<$ 1.0 mm, and MSD $<$ 1.41 mm. Conclusion. The proposed combination of a DL-based and iodine contrast-enhanced model proved as a viable method to vastly improve efficiency of small animal CBCT image-guided RT preclinical trials.

physics.med-ph

DRKF: Distilled Rotated Kernel Fusion for Efficient Rotation Invariant Descriptors in Local Feature Matching

The performance of local feature descriptors degrades in the presence of large rotation variations. To address this issue, we present an efficient approach to learning rotation invariant descriptors. Specifically, we propose Rotated Kernel Fusion (RKF) which imposes rotations on the convolution kernel to improve the inherent nature of CNN. Since RKF can be processed by the subsequent re-parameterization, no extra computational costs will be introduced in the inference stage. Moreover, we present Multi-oriented Feature Aggregation (MOFA) which aggregates features extracted from multiple rotated versions of the input image and can provide auxiliary knowledge for the training of RKF by leveraging the distillation strategy. We refer to the distilled RKF model as DRKF. Besides the evaluation on a rotation-augmented version of the public dataset HPatches, we also contribute a new dataset named DiverseBEV which is collected during the drone's flight and consists of bird's eye view images with large viewpoint changes and camera rotations. Extensive experiments show that our method can outperform other state-of-the-art techniques when exposed to large rotation variations.

cs.CV

Mechanism and Model of a Soft Robot for Head Stabilization in Cancer Radiation Therapy

We present a parallel robot mechanism and the constitutive laws that govern the deformation of its constituent soft actuators. Our ultimate goal is the real-time motion-correction of a patient's head deviation from a target pose where the soft actuators control the position of the patient's cranial region on a treatment machine. We describe the mechanism, derive the stress-strain constitutive laws for the individual actuators and the inverse kinematics that prescribes a given deformation, and then present simulation results that validate our mathematical formulation. Our results demonstrate deformations consistent with our radially symmetric displacement formulation under a finite elastic deformation framework.

cs.RO

Improving the efficiency of small animal 3D printed compensator IMRT with beamlet intensity total variation regularization

Purpose: There is growing interest in the use of modern 3D printing technology to implement intensity-modulated radiation therapy (IMRT) on the preclinical scale which is analogous to clinical IMRT. However, current 3D-printed IMRT methods suffer from complex modulation patterns leading to long delivery times, excess filament usage, and inaccurate compensator fabrication. In this work, we have developed a total variation regularization (TVR) approach to address these issues. Methods: TVR-IMRT, a technique designed to minimize the intensity difference between neighboring beamlets, was used to optimize the beamlet intensity map, which was then converted to corresponding compensator thicknesses in copper-doped PLA filament. IMRT and TVR-IMRT plans using five beams were generated to treat a mouse heart while sparing lung tissue. The individual field doses and composite dose were delivered to film and compared to the corresponding planned doses using gamma analysis. Results: TVR-IMRT reduced the total variation of both the beamlet intensities and compensator thicknesses by around 50% when compared to standard 3D printed compensator IMRT. The total mass of compensator material consumed and radiation beam-on time were reduced by 20-30%, while DVHs remained comparable. Gamma analysis passing rate with 3%/0.3mm criterion was 89.07% for IMRT and 95.37% for TVR-IMRT. Conclusion: TVR can be applied to small animal IMRT beamlet intensities in order to produce fluence maps and subsequent 3D-printed compensator patterns with less total variation, simplifying 3D printing and reducing the amount of filament required. The TVR-IMRT plan required less beam-on time while maintaining the dose conformity when compared to a traditional IMRT plan.

physics.med-ph

A conceptual study on real-time adaptive radiation therapy optimization through ultra-fast beamlet control

A central problem in the field of radiation therapy (RT) is how to optimally deliver dose to a patient in a way that fully accounts for anatomical position changes over time. As current RT is a static process, where beam intensities are calculated before the start of treatment, anatomical deviations can result in poor dose conformity. To overcome these limitations, we present a simulation study on a fully dynamic real-time adaptive radiation therapy (RT-ART) optimization approach that uses ultra-fast beamlet control to dynamically adapt to patient motion in real-time. A virtual RT-ART machine was simulated with a rapidly rotating linear accelerator (LINAC) source (60 RPM) and a binary 1D multi-leaf collimator (MLC) operating at 100 Hz. If the real-time tracked target motion exceeded a predefined threshold, a time dependent objective function was solved using fast optimization methods to calculate new beamlet intensities that were then delivered to the patient. To evaluate the approach, system response was analyzed for patient derived continuous drift, step-like, and periodic intra-fractional motion. For each motion type investigated, the RT-ART method was compared against the ideal case with no patient motion (static case) as well as to the case without the use RT-ART. In all cases, isodose lines and dose-volume-histograms (DVH) showed that RT-ART plan quality was approximately the same as the static case, and considerably better than the no RT-ART case. The RT-ART optimization framework has the potential to optimally deliver dose to a patient in a way that fully accounts for anatomical changes due to motion. With continued advances in real-time patient motion tracking and fast computational processes, there is significant potential for the RT-ART optimization process to be realized on next generation RT machines.

physics.med-ph

Trajectory planning optimization for real-time 6DOF robotic patient motion compensation

We present for the first time a general 6DoF trajectory planning method that can be used in real-time image guided radiation therapy procedures for robotic stabilization of dynamically moving tumor targets. As the radiation beam is always on during the motion compensation process, it is mandatory that the 6D correction trajectory is optimal both spatially and temporally in order to maximize radiation to the tumor and minimize unintentional irradiation of healthy tissues. Unlike prior works, which relied on motion control approaches as PID or other controllers, this work presents the concept of motion planning, where all potential 6D trajectories are searched using ultrafast optimization methods and the best trajectory is chosen. As the method formulates the problem as an objective function to be solved, it allows high flexibility in that users can optimize various performance requirements such as mechanical robot limits, patient velocities, or other aspects that must operate within certain limits in order to ensure a safe medical process.

physics.med-ph

On Differential Geometric Approach to Nonlinear Systems Affine in Control

The note focuses on the differential geometric approach to the study of nonlinear systems that are affine in control. We first develop normal forms for nonlinear system affine in control. Based on these normal forms, we then address the problems of global stabilization, semi-global stabilization and disturbance attenuation.

math.DS