Searcharxiv⌕ Search

arXiv subjects

Jianxiong Wu

Publications and source records attributed to Jianxiong Wu.

8 recordsLinked to original sources

LoRango: It Takes Two LoRAs to Unlock Hidden Behaviors in Diffusion Models

Users commonly combine multiple Low-Rank Adaptation (LoRA) adapters to personalize images with different subjects, styles, and visual attributes. Yet inspecting adapters individually does not establish the safety of their composition. We identify and characterize a pair-conditioned attack in text-to-image diffusion: individually useful and benign-appearing adapters redirect image generation when co-loaded with a specifically matched partner, whose identity serves as the trigger. We introduce LoRango to realize this attack through complementary Signature and Payload adapters. The Signature writes a pair-specific code into intermediate carrier representations, while the Payload uses code-selective responses and opposing signal/reference branches. These branches approximately cancel for standalone adapters and mismatched pairs; matched code-reader alignment breaks cancellation within native GEGLU blocks and releases the programmed action. Both adapters are exported as ordinary static LoRA files compatible with standard loaders, requiring no prompt trigger or base-pipeline modification. LoRango achieves matched-pair attack success rates of 97.9\% on SD v1.5 and 98.7\% on SDXL, compared with 2.8--4.6\% when implanted adapters are loaded individually. Further experiments evaluate pair selectivity, standalone fidelity, robustness to deployment variations, and applicability across denoiser architectures. These findings show that individual-adapter inspection is insufficient to assess the security of multi-LoRA personalization and motivate auditing adapter compositions.

cs.CR↗

CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies

Flow-based vision-language-action (VLA) policies offer strong expressivity for action generation, but suffer from a fundamental inefficiency: multi-step inference is required to recover action structure from uninformative Gaussian noise, leading to a poor efficiency-quality trade-off under real-time constraints. We address this issue by rethinking the role of the starting point in generative action modeling. Instead of shortening the sampling trajectory, we propose CF-VLA, a coarse-to-fine two-stage formulation that restructures action generation into a coarse initialization step that constructs an action-aware starting point, followed by a single-step local refinement that corrects residual errors. Concretely, the coarse stage learns a conditional posterior over endpoint velocity to transform Gaussian noise into a structured initialization, while the fine stage performs a fixed-time refinement from this initialization. To stabilize training, we introduce a stepwise strategy that first learns a controlled coarse predictor and then performs joint optimization. Experiments on CALVIN and LIBERO show that our method establishes a strong efficiency-performance frontier under low-NFE (Number of Function Evaluations) regimes: it consistently outperforms existing NFE=2 methods, matches or surpasses the NFE=10 $π_{0.5}$ baseline on several metrics, reduces action sampling latency by 75.4%, and achieves the best average real-robot success rate of 83.0%, outperforming MIP by 19.5 points and $π_{0.5}$ by 4.0 points. These results suggest that structured, coarse-to-fine generation enables both strong performance and efficient inference. Our code is available at https://github.com/EmbodiedAI-RoboTron/CF-VLA.

cs.CV↗

A Novel Evaluation Benchmark for Medical LLMs: Illuminating Safety and Effectiveness in Clinical Domains

Large language models (LLMs) hold promise in clinical decision support but face major challenges in safety evaluation and effectiveness validation. We developed the Clinical Safety-Effectiveness Dual-Track Benchmark (CSEDB), a multidimensional framework built on clinical expert consensus, encompassing 30 criteria covering critical areas like critical illness recognition, guideline adherence, and medication safety, with weighted consequence measures. Thirty-two specialist physicians developed and reviewed 2,069 open-ended Q&A items aligned with these criteria, spanning 26 clinical departments to simulate real-world scenarios. Benchmark testing of six LLMs revealed moderate overall performance (average total score 57.2%, safety 54.7%, effectiveness 62.3%), with a significant 13.3% performance drop in high-risk scenarios (p < 0.0001). Domain-specific medical LLMs showed consistent performance advantages over general-purpose models, with relatively higher top scores in safety (0.912) and effectiveness (0.861). The findings of this study not only provide a standardized metric for evaluating the clinical application of medical LLMs, facilitating comparative analyses, risk exposure identification, and improvement directions across different scenarios, but also hold the potential to promote safer and more effective deployment of large language models in healthcare environments.

cs.CL↗

Semiquantum key distribution using initial states in only one basis without the classical user measuring

From the perspective of resource theory, it is interesting to achieve the same quantum task using as few quantum resources as possible. Semiquantum key distribution (SQKD), which allows a quantum user to share a confidential key with a classical user who prepares and operates qubits in only one basis, is an important example for studying this issue. To further limit the quantum resources used by users, in this paper, we constructed the first SQKD protocol which restricts the quantum user to prepare quantum states in only one basis and removes the classical user's measurement capability. Furthermore, we prove that the constructed protocol is unconditionally secure by deriving a key rate expression of the error rate in the asymptotic scenario. The work of this paper provides inspiration for achieving quantum superiority with minimal quantum resources.

quant-ph↗

A Teacher-Student Framework for Semi-supervised Medical Image Segmentation From Mixed Supervision

Standard segmentation of medical images based on full-supervised convolutional networks demands accurate dense annotations. Such learning framework is built on laborious manual annotation with restrict demands for expertise, leading to insufficient high-quality labels. To overcome such limitation and exploit massive weakly labeled data, we relaxed the rigid labeling requirement and developed a semi-supervised learning framework based on a teacher-student fashion for organ and lesion segmentation with partial dense-labeled supervision and supplementary loose bounding-box supervision which are easier to acquire. Observing the geometrical relation of an organ and its inner lesions in most cases, we propose a hierarchical organ-to-lesion (O2L) attention module in a teacher segmentor to produce pseudo-labels. Then a student segmentor is trained with combinations of manual-labeled and pseudo-labeled annotations. We further proposed a localization branch realized via an aggregation of high-level features in a deep decoder to predict locations of organ and lesion, which enriches student segmentor with precise localization information. We validated each design in our model on LiTS challenge datasets by ablation study and showed its state-of-the-art performance compared with recent methods. We show our model is robust to the quality of bounding box and achieves comparable performance compared with full-supervised learning methods.

cs.CV↗

Discrete solitons in waveguide arrays with long-range linearly coupled effect

We study the influences to the discrete soliton (DS) by introducing linearly long-range nonlocal interactions, which give rise to the off-diagonal elements of the linearly coupled matrix in the discrete nonlinear schrodinger equation to be filled by non-zero terms. Theoretical analysis and numerical simulations find that the DS under this circumstance can exhibit strong digital effects: the fundamental DS is a narrow one, which occupies nearly only one waveguide, the dipole and double-monopole solitons, which occupy two waveguides, can be found in self-focusing and -defocusing nonlinearities, respectively. Stable flat-top solitons and their stagger counterparts, which occupy a controllable number of waveguides, can also be obtained through this system. Such digital properties may give rise to additional data processing applications and have potential in fabricating digital optical devices in all-optical networks.

physics.optics↗

Switch between the types of the symmetry breaking bifurcation in optically induced photorefractive rotational double-well potential

We study the possibility of switching the types of symmetry breaking bifurcation (SBB) in the cylinder shell waveguide with helical double-well potential along propagation direction. This model is described by the one-dimensional nonlinear Schrödinger (NLS) equation. The symmetry- and antisymmetry-breakings can be caused by increasing the applied voltage onto the waveguide in the self-focusing and -defocusing cases, respectively. In the self-focusing case, the type of SBB can be switched from supercritical to subcritical. While in the self-defocusing case, the type of SBB can not be switched because only one type of SBB is found.

physics.optics↗

Quasi-compactons in inverted nonlinear photonic crystals

We study large-amplitude one-dimensional solitary waves in photonic crystals featuring competition between linear and nonlinear lattices, with minima of the linear potential coinciding with maxima of the nonlinear pseudopotential, and vice versa (inverted nonlinear photonic crystals, INPhCs), in the case of the saturable self-focusing nonlinearity. Such crystals were recently fabricated using a mixture of SU-8 and Rhodamine-B optical materials. By means of numerical methods and analytical approximations, we find that large-amplitude solitons are broad sharply localized stable pulses (quasi-compactons, QCs). With the increase of the totalpower, P, the QC's centroid performs multiple switchings between minima and maxima of the linear potential. Unlike cubic INPhCs, the large-amplitude solitons are mobile in the medium with the saturable nonlinearity. The threshold value of the kick necessary to set the soliton in motion is found as a function of P. Collisions between moving QCs are considered too.

physics.optics↗