SearcharxivSearch

arXiv subjects

Wenbin He

Publications and source records attributed to Wenbin He.

At least 19 recordsLinked to original sources

Understanding Software Defect Prediction: A Large-scale Empirical Study Across Uncertainty Quantification and Performance Evaluation

Software defect prediction (SDP) classifiers produce probabilities used for inspection prioritization, threshold tuning, and risk communication. Probability-based uncertainty quantification (UQ) characterizes prediction confidence, but whether common UQ metrics reliably indicate performance and calibration remains unclear. We conducted a large-scale empirical study of probability-based UQ for SDP. We evaluated five UQ metrics, six performance metrics, and three calibration metrics for 16 representative classifiers. We analyzed these relationships under two prediction settings: within-project defect prediction (WPDP), using 36 benchmark datasets, and cross-project defect prediction (CPDP), using 32 feature-compatible datasets. Results showed that UQ was highly context-dependent. Under WPDP, UQ correlated more consistently with false positive rate and AUC than with MCC, F1 score, and other metrics; these correlations also varied across classifier categories and dataset collections. Performance and calibration were related but not interchangeable; classifiers with strong discrimination could still exhibit large calibration error. Under CPDP, several UQ-performance and UQ-calibration correlations weakened or reversed, indicating that uncertainty signals do not reliably transfer across projects. Thus, UQ should be evaluated against specific performance objectives. Calibration should be assessed independently using multiple metrics. Transferred probabilities should be revalidated before guiding quality-assurance decisions.

cs.SE

Low-threshold efficient N${_2^+}$ lasing driven by sub-cycle soliton dynamics in a hollow waveguide

The phenomenon of N${_2^+}$ lasing, observed in femtosecond-laser filamentation, attract considerable interests in recent several years, with great application potentials in fields of remote sensing and ultrafast spectroscopy. Efficient N${_2^+}$ lasing at relatively-low pump energies and with high beam quality, while being highly-demanded for applications, remains, however, quite challenging in practical experiments. Here, we demonstrate a new route of generating low-threshold N${_2^+}$ lasing with unprecedently-high efficiency, which is enabled by soliton dynamics in a gas-filled hollow-tapered-capillary system. High-order-soliton compression of a 12-fs, 10-${\mu}$J-level pump pulse forms a sub-cycle asymmetric transient that tunnel-ionizes N${_2}$ to N${_2^+}$ and, through direct, single-photon resonant excitation, creates population inversion between the ground state ${X^2\Sigma_g^+}$ and the excited state ${B^2\Sigma_u^+}$${-}$a dynamic process distinct from the widely adopted three-state coupling picture${-}$and remarkably at unexpectedly low pump energy. In the experiments, we obtained 100-nJ-level N${_2^+}$ lasing pulses at 391 nm with conversion efficiencies up to 3.3$\times$10$^{-3}$, at pump energies of less than 50 ${\mu}$J. These results represent improvement of more than one orders of magnitude in both generation efficiency and lasing threshold, compared with prevailing filamentation-based schemes. Our study bridges two generally-disparate fields (sub-cycle soliton dynamics and N${_2^+}$ lasing), and paves the way for narrow-band, high-beam-quality lasing pulses that may find wide applications in advanced spectroscopy and nonlinear pump-probe experiments.

physics.optics

Scaling Expert Feedback with Reflective Edit Propagation in Compositional Knowledge Bases

Domain-specific knowledge bases (KBs) encode vertical expertise and proprietary information that organizations depend on, but curating them at scale is a persistent challenge. Although Large Language Models (LLMs) can draft initial entries efficiently, technical accuracy still requires human expert validation, and reviewing entries one by one at scale is impractical. We present Reflective Agent for Identifier Dictionary (RAID), a novel system that transforms individual expert edits into systematic knowledge updates. Unlike traditional "correct-and-save" paradigms, RAID utilizes a reflective agent to infer the underlying semantic intent behind a single expert edit and propagates that correction across the entire KB through a three-step architecture: Intent Inference, Reflection-based Planning, and User Controlled Execution. We evaluated the reflection and propagation performance on a public dataset and conducted a user study with subject matter experts with proprietary data. The evaluation shows RAID's technical feasibility in capturing expert intent and its potential to scale specialized expertise across industrial knowledge bases.

cs.HC

Enhancement of vacuum-ultraviolet dispersive-wave emission using gas-filled tapered hollow-core fibers

The recent breakthroughs in laser-driving 229Th nuclear transition have created an urgent demand for coherent vacuum-ultraviolet (VUV) sources delivering high spectral brightness at the critical 148.38 nm isomer energy. However, generating sufficient photon flux to overcome the low nuclear excitation probability remains a challenge for compact setups. While resonant dispersive wave emission in gas-filled hollow-core fibers offers a promising route, standard capillaries face a fundamental trade-off: maximizing input coupling requires large core diameters, whereas efficient nonlinear VUV conversion demands the high intensities using small cores. Here, we resolve this conflict using a gas-filled tapered capillary fiber. This architecture utilizes a longitudinally decreasing core diameter to combine a large input aperture with adiabatic field concentration, thereby continuously enhancing the nonlinear interaction. Experimentally, we demonstrate a widely tunable source (135-240 nm) that achieves a twofold efficiency enhancement specifically at the 148.38 nm wavelength compared to uniform geometries. By providing a scalable route to high-flux VUV generation, this work establishes a critical tabletop tool for advancing solid-state nuclear clocks and time-resolved spectroscopy.

physics.optics

Deep-ultraviolet Cherenkov radiation in all-normal-dispersion waveguide enabled by spatial-temporal dynamics

Nonlinear propagation of ultrashort pulses in multi-mode waveguides, featuring complex spatial-temporal dynamics, provides new degrees of freedom in the fields of nonlinear optics and ultrafast lasers. Here, we demonstrate a new scheme of ultraviolet Cherenkov (dispersive-wave) radiation in a gas-filled capillary with unprecedently-high pulse energy, enabled by spatial-temporal dynamics. We found that mJ-level, 40-fs pulses, launched into a large-core capillary filled with high-pressure noble gas, would experience self-phase-modulation and self-steepening effects in this normal-dispersion waveguide, leading to high-intensity shock wave generation and asymmetric spectral broadening. Spatial-temporal dynamics, stemming from strong nonlinear inter-mode coupling, causes spatial shrink and temporal deceleration of the pulse which dramatically alter the capillary dispersion landscape. As a result, a phase-matching point can be created in the ultraviolet, giving rise to the radiation of multi-mode dispersive waves with 100-{\mu}J-level pulse energies and few-fs pulse widths. Our findings inspire new insights into multi-mode nonlinear optics, and the demonstrated high-energy ultraviolet light source with broadband tunability and compact set-up configuration, may find a few applications in time-resolved spectroscopy, ultrafast electronics and femtosecond chemistry.

physics.optics

Hierarchical self-organization of highly-ordered granular ensemble of optical solitons through collective motions

Self-organizations of ordered patterns in far-from-equilibrium many-body systems host fundamental importance in many disciplines. Meanwhile, complex systems often feature hierarchical structures with distinct scales for different layers, enabling high-level effective dynamics without exhaustive tracking of all possible degrees of freedoms. In this work, we report a study of the self-organization dynamics of highly-ordered soliton ensembles in a high-harmonic mode-locked fiber lasers through collective motions driven by nonlocal optomechanical interactions and local collisions, which exhibit a series of universal characteristics reminiscent of phase transitions. Moreover, the multi-soliton laser-field can be coarsely grained as a granular ensemble of limit-cycle oscillators with simple interaction rules derived from fine-scale physics. The self-organization of the multitude of solitons in the mode-locked laser cavity can then be mapped into a low-dimensional dynamic model that essentially reproduced the emergent process. Our work affords a conceptual framework for understanding the complex structure formation in nonlinear laser systems, and may help to design ultrafast lasers by exploiting universal principles of collective motions.

physics.optics

Rethinking Agentic Workflows: Evaluating Inference-Based Test-Time Scaling Strategies in Text2SQL Tasks

Large language models (LLMs) are increasingly powering Text-to-SQL (Text2SQL) systems, enabling non-expert users to query industrial databases using natural language. While test-time scaling strategies have shown promise in LLM-based solutions, their effectiveness in real-world applications, especially with the latest reasoning models, remains uncertain. In this work, we benchmark six lightweight, industry-oriented test-time scaling strategies and four LLMs, including two reasoning models, evaluating their performance on the BIRD Mini-Dev benchmark. Beyond standard accuracy metrics, we also report inference latency and token consumption, providing insights relevant for practical system deployment. Our findings reveal that Divide-and-Conquer prompting and few-shot demonstrations consistently enhance performance for both general-purpose and reasoning-focused LLMs. However, introducing additional workflow steps yields mixed results, and base model selection plays a critical role. This work sheds light on the practical trade-offs between accuracy, efficiency, and complexity when deploying Text2SQL systems.

cs.CL

VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels

The advances in multi-modal foundation models (FMs) (e.g., CLIP and LLaVA) have facilitated the auto-labeling of large-scale datasets, enhancing model performance in challenging downstream tasks such as open-vocabulary object detection and segmentation. However, the quality of FM-generated labels is less studied as existing approaches focus more on data quantity over quality. This is because validating large volumes of data without ground truth presents a considerable challenge in practice. Existing methods typically rely on limited metrics to identify problematic data, lacking a comprehensive perspective, or apply human validation to only a small data fraction, failing to address the full spectrum of potential issues. To overcome these challenges, we introduce VISTA, a visual analytics framework that improves data quality to enhance the performance of multi-modal models. Targeting the complex and demanding domain of open-vocabulary image segmentation, VISTA integrates multi-phased data validation strategies with human expertise, enabling humans to identify, understand, and correct hidden issues within FM-generated labels. Through detailed use cases on two benchmark datasets and expert reviews, we demonstrate VISTA's effectiveness from both quantitative and qualitative perspectives.

cs.CV

ProSAM: Enhancing the Robustness of SAM-based Visual Reference Segmentation with Probabilistic Prompts

The recent advancements in large foundation models have driven the success of open-set image segmentation, a task focused on segmenting objects beyond predefined categories. Among various prompt types (such as points, boxes, texts, and visual references), visual reference segmentation stands out for its unique flexibility and strong zero-shot capabilities. Recently, several SAM-based methods have made notable progress in this task by automatically generating prompts to guide SAM. However, these methods often generate prompts at boundaries of target regions due to suboptimal prompt encoder, which results in instability and reduced robustness. In this work, we introduce ProSAM, a simple but effective method to address the stability challenges we identified in existing SAM-based visual reference segmentation approaches. By learning a variational prompt encoder to predict multivariate prompt distributions, ProSAM avoids generating prompts that lie in unstable regions, overcoming the instability caused by less robust prompts. Our approach consistently surpasses state-of-the-art methods on the Pascal-5$^i$ and COCO-20$^i$ datasets, providing a more robust solution for visual reference segmentation.

cs.CV

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads

Vision foundation models (VFMs) have demonstrated remarkable performance across a wide range of downstream tasks. While several VFM adapters have shown promising results by leveraging the prior knowledge of VFMs, we identify two inefficiencies in these approaches. First, the interaction between convolutional neural network (CNN) and VFM backbone triggers early layer gradient backpropagation. Second, existing methods require tuning all components, adding complexity. Besides, these adapters alter VFM features, underutilizing the prior knowledge. To tackle these challenges, we propose a new approach called ViT-Split, based on a key observation: the layers of several VFMs, like DINOv2, can be divided into two distinct components: an extractor for learning low-level features and an adapter for learning task-specific features. Leveraging this insight, we eliminate the CNN branch and introduce two heads, task head and prior head, to the frozen VFM. The task head is designed to learn task-specific features, mitigating the early gradient propagation issue. The prior head is used to leverage the multi-scale prior features from the frozen VFM, reducing tuning parameters and overfitting. Extensive experiments on various tasks (e.g., segmentation, detection, depth estimation, and visual question answering) validate the effectiveness and efficiency of ViT-Split. Specifically, ViT-Split reduces training time up to $4\times$ while achieving comparable or even better results on ADE20K, compared to other VFM adapters.

cs.CV

DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models

The recent explosive interest in the reasoning capabilities of large language models, such as DeepSeek-R1, has demonstrated remarkable success through reinforcement learning-based fine-tuning frameworks, exemplified by methods like Group Relative Policy Optimization (GRPO). However, such reasoning abilities remain underexplored and notably absent in vision foundation models, including representation models like the DINO series. In this work, we propose \textbf{DINO-R1}, the first such attempt to incentivize visual in-context reasoning capabilities of vision foundation models using reinforcement learning. Specifically, DINO-R1 introduces \textbf{Group Relative Query Optimization (GRQO)}, a novel reinforcement-style training strategy explicitly designed for query-based representation models, which computes query-level rewards based on group-normalized alignment quality. We also apply KL-regularization to stabilize the objectness distribution to reduce the training instability. This joint optimization enables dense and expressive supervision across queries while mitigating overfitting and distributional drift. Building upon Grounding-DINO, we train a series of DINO-R1 family models that integrate a visual prompt encoder and a visual-guided query selection mechanism. Extensive experiments on COCO, LVIS, and ODinW demonstrate that DINO-R1 significantly outperforms supervised fine-tuning baselines, achieving strong generalization in both open-vocabulary and closed-set visual prompting scenarios.

cs.CV

Reducing Averaging Time in Dual-comb Spectroscopy via Phase-Patterned Higher-Repetition-Rate Pulses

Dual-comb spectroscopy (DCS) is a powerful Fourier-transform spectroscopic technique that provides high-speed, high-resolution, and broadband measurements without moving parts. However, the high peak power of mode-locked pulses limits the photodetector's dynamic range, resulting in a low signal-to-noise ratio (SNR) per acquisition. While coherent averaging can improve SNR, it sacrifices temporal resolution and demands stringent system stability. Here, we introduce a novel concept to enhance SNR by using phase-patterned higher-repetition-rate combs. We reinterpret the self-imaging process of comb spectrum from a new perspective on mode interference among sub-pulse trains As a proof-of-concept, we densified two 250-MHz frequency combs to 12.5-MHz mode spacings via phase modulation and performed DCS on an $\mathrm{H^{13}C^{14}N}$ gas cell, and compared the results with an emulated conventional 12.5-MHz DCS, demonstrating a 17-fold increase in mode amplitude. This concept is expected to be combined with ultra-high repetition rate combs, such as microcombs, and thereby deployed in practical applications that typically require spectral sampling spacings from hundreds of MHz to GHz range.

physics.optics

InterChat: Enhancing Generative Visual Analytics using Multimodal Interactions

The rise of Large Language Models (LLMs) and generative visual analytics systems has transformed data-driven insights, yet significant challenges persist in accurately interpreting users' analytical and interaction intents. While language inputs offer flexibility, they often lack precision, making the expression of complex intents inefficient, error-prone, and time-intensive. To address these limitations, we investigate the design space of multimodal interactions for generative visual analytics through a literature review and pilot brainstorming sessions. Building on these insights, we introduce a highly extensible workflow that integrates multiple LLM agents for intent inference and visualization generation. We develop InterChat, a generative visual analytics system that combines direct manipulation of visual elements with natural language inputs. This integration enables precise intent communication and supports progressive, visually driven exploratory data analyses. By employing effective prompt engineering, and contextual interaction linking, alongside intuitive visualization and interaction designs, InterChat bridges the gap between user interactions and LLM-driven visualizations, enhancing both interpretability and usability. Extensive evaluations, including two usage scenarios, a user study, and expert feedback, demonstrate the effectiveness of InterChat. Results show significant improvements in the accuracy and efficiency of handling complex visual analytics tasks, highlighting the potential of multimodal interactions to redefine user engagement and analytical depth in generative visual analytics.

cs.HC

Flexible delivery of high-power picosecond laser in purely-single optical mode of anti-resonant hollow-core fiber for micromachining

We present the flexible delivery of picosecond laser pulses with up to 20 W average power over a 3-m-long sample of anti-resonant hollow-core fiber (AR-HCF) for laser micromachining applications. Our experiments highlight the importance of optical mode purity of the AR-HCF for the manufacturing precision. We demonstrate that compared with an AR-HCF sample with a capillary to core (d/D) ratio of ~0.5, the AR-HCF with a d/D ratio of ~0.68 exhibits better capability of high-order-mode suppression, giving rise to improved micromachining quality. Moreover, the AR-HCF delivery system exhibits better pointing stability and set-up flexibility than the free-space beam delivery system. These results pave the way to practical applications of AR-HCF in developing advanced equipment for ultrafast laser micromachining.

physics.optics

Tunable ultraviolet dispersive-wave emission driven directly by 40-fs Ti: sapphire laser pulses in hollow capillary fiber

We demonstrate that by using 1-m-long gas-filled hollow capillary fiber (HCF) with a core diameter of 100 {\mu}m, tunable ultraviolet (UV) dispersive-wave (DW) pulses can be generated in a compact, single-stage set-up driven directly by 40-fs Ti: sapphire laser pulses. By adjusting the gas type and pressure inside the HCF, the central wavelength of the UV DW can be continuously tuned from 185 nm to ~450 nm. In the experiment, we found that for longer-wavelength (from ~320 to ~450 nm) DW generation, Raman-active gas filled in the HCF can efficiently suppress the pulse splitting effect of the high-order soliton due to the Raman-induced pulse energy dissipation, leading to the high-quality DW generation at these wavelengths with smooth, single-peak spectra. These results provide some useful insights for designing compact, wavelength-tunable ultrafast UV light sources with microjoule-level pulse energies.

physics.optics

InFiConD: Interactive No-code Fine-tuning with Concept-based Knowledge Distillation

The emergence of large-scale pre-trained models has heightened their application in various downstream tasks, yet deployment is a challenge in environments with limited computational resources. Knowledge distillation has emerged as a solution in such scenarios, whereby knowledge from large teacher models is transferred into smaller student' models, but this is a non-trivial process that traditionally requires technical expertise in AI/ML. To address these challenges, this paper presents InFiConD, a novel framework that leverages visual concepts to implement the knowledge distillation process and enable subsequent no-code fine-tuning of student models. We develop a novel knowledge distillation pipeline based on extracting text-aligned visual concepts from a concept corpus using multimodal models, and construct highly interpretable linear student models based on visual concepts that mimic a teacher model in a response-based manner. InFiConD's interface allows users to interactively fine-tune the student model by manipulating concept influences directly in the user interface. We validate InFiConD via a robust usage scenario and user study. Our findings indicate that InFiConD's human-in-the-loop and visualization-driven approach enables users to effectively create and analyze student models, understand how knowledge is transferred, and efficiently perform fine-tuning operations. We discuss how this work highlights the potential of interactive and visual methods in making knowledge distillation and subsequent no-code fine-tuning more accessible and adaptable to a wider range of users with domain-specific demands.

cs.LG

Retiming dynamics of harmonically modelocked laser solitons in a self-driven optomechanical lattice

Harmonic mode-locking, realized actively or passively, is an effective technique for increasing the repetition rate of lasers, with important applications in optical sampling, laser micro-machining and frequency metrology. It is critically important to understand how a harmonically mode-locked pulse train responds to external perturbations and noise, so as to make sure that it is stable and resistant to noise. Here, in a series of carefully designed experiments, we elucidate the retiming dynamics of laser pulses generated in a soliton fiber laser harmonically mode-locked at ~2 GHz to the acoustic resonance in a photonic crystal fiber (PCF) core. We characterize the self-driven optomechanical lattice along the PCF using a homodyne set-up, and reveal that each soliton undergoes damped oscillatory retiming within its trapping potential after an abrupt perturbation. In addition we show, through statistical analysis of the intra-cavity pulse spacing, how the trapping potentials are effective for suppressing timing jitter. The experimental results are well described using a dynamic model including dissipation, which provides valuable insight into the stability and noise performance of optomechanically mode-locked laser systems, and may also be useful for studying complex inter-soliton interactions.

physics.optics

USE: Universal Segment Embeddings for Open-Vocabulary Image Segmentation

The open-vocabulary image segmentation task involves partitioning images into semantically meaningful segments and classifying them with flexible text-defined categories. The recent vision-based foundation models such as the Segment Anything Model (SAM) have shown superior performance in generating class-agnostic image segments. The main challenge in open-vocabulary image segmentation now lies in accurately classifying these segments into text-defined categories. In this paper, we introduce the Universal Segment Embedding (USE) framework to address this challenge. This framework is comprised of two key components: 1) a data pipeline designed to efficiently curate a large amount of segment-text pairs at various granularities, and 2) a universal segment embedding model that enables precise segment classification into a vast range of text-defined categories. The USE model can not only help open-vocabulary image segmentation but also facilitate other downstream tasks (e.g., querying and ranking). Through comprehensive experimental studies on semantic segmentation and part segmentation benchmarks, we demonstrate that the USE framework outperforms state-of-the-art open-vocabulary segmentation methods.

cs.CV