SearcharxivSearch

arXiv subjects

Po-Han Huang

Publications and source records attributed to Po-Han Huang.

9 recordsLinked to original sources

DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection

Vision foundation models have enabled strong training-free anomaly detection (AD). However, most existing approaches rely primarily on independent local patch features, leaving the global contextual information encoded by Vision Transformers (ViTs) underexploited. In this work, we identify the dual characteristics of the ViT [CLS] token: its embedding provides anomaly-invariant global semantic representation, while its attention maps implicitly highlight spatially abnormal regions. Building on this observation, we propose a fully automated AD framework leveraging global context to remove manual tunings. Our framework introduces (1) an automatic augmentation selection strategy driven by [CLS]-level semantic consistency, and (2) an attention-guided feature reweighting mechanism that dynamically adjusts patch contributions according to [CLS] attention saliency. By integrating these components over multi-level features, our method achieves stable anomaly scoring and precise localization without training or parameter tuning. Under the one-shot setting, it achieves Image-AUC scores of 97.7%, 93.2%, and 84.5% on MVTec-AD, VisA, and Real-IAD. Using a single fixed configuration across categories, backbones, and datasets, the method establishes a new state-of-the-art for plug-and-play, training-free anomaly detection while maintaining strong robustness and practical scalability.

cs.CV

PatchEAD: Unifying Industrial Visual Prompting Frameworks for Patch-Exclusive Anomaly Detection

Industrial anomaly detection is increasingly relying on foundation models, aiming for strong out-of-distribution generalization and rapid adaptation in real-world deployments. Notably, past studies have primarily focused on textual prompt tuning, leaving the intrinsic visual counterpart fragmented into processing steps specific to each foundation model. We aim to address this limitation by proposing a unified patch-focused framework, Patch-Exclusive Anomaly Detection (PatchEAD), enabling training-free anomaly detection that is compatible with diverse foundation models. The framework constructs visual prompting techniques, including an alignment module and foreground masking. Our experiments show superior few-shot and batch zero-shot performance compared to prior work, despite the absence of textual features. Our study further examines how backbone structure and pretrained characteristics affect patch-similarity robustness, providing actionable guidance for selecting and configuring foundation models for real-world visual inspection. These results confirm that a well-unified patch-only framework can enable quick, calibration-light deployment without the need for carefully engineered textual prompts.

cs.CV

3D printing of hierarchical structures made of inorganic silicon-rich glass featuring self-forming nanogratings

Hierarchical structures are abundant in nature, such as in the superhydrophobic surfaces of lotus leaves and the structural coloration of butterfly wings. They consist of ordered features across multiple size scales, and their unique properties have attracted enormous interest in wide-ranging fields, including energy storage, nanofluidics, and nanophotonics. Femtosecond lasers, capable of inducing various material modifications, have shown promise for manufacturing tailored hierarchical structures. However, existing methods such as multiphoton lithography and 3D printing using nanoparticle-filled inks typically involve polymers and suffer from high process complexity. Here, we demonstrate 3D printing of hierarchical structures in inorganic silicon-rich glass featuring self-forming nanogratings. This approach takes advantage of our finding that femtosecond laser pulses can induce simultaneous multiphoton crosslinking and self-formation of nanogratings in hydrogen silsesquioxane (HSQ). The 3D printing process combines the 3D patterning capability of multiphoton lithography and the efficient generation of periodic structures by the self-formation of nanogratings. We 3D-printed micro-supercapacitors with large surface areas and a remarkable areal capacitance of 1 mF/cm^2 at an ultrahigh scan rate of 50 V/s, thereby demonstrating the utility of our 3D printing approach for device applications in emerging fields such as energy storage.

physics.app-ph

Improving Limited Supervised Foot Ulcer Segmentation Using Cross-Domain Augmentation

Diabetic foot ulcers pose health risks, including higher morbidity, mortality, and amputation rates. Monitoring wound areas is crucial for proper care, but manual segmentation is subjective due to complex wound features and background variation. Expert annotations are costly and time-intensive, thus hampering large dataset creation. Existing segmentation models relying on extensive annotations are impractical in real-world scenarios with limited annotated data. In this paper, we propose a cross-domain augmentation method named TransMix that combines Augmented Global Pre-training AGP and Localized CutMix Fine-tuning LCF to enrich wound segmentation data for model learning. TransMix can effectively improve the foot ulcer segmentation model training by leveraging other dermatology datasets not on ulcer skins or wounds. AGP effectively increases the overall image variability, while LCF increases the diversity of wound regions. Experimental results show that TransMix increases the variability of wound regions and substantially improves the Dice score for models trained with only 40 annotated images under various proportions.

cs.CV

Graphene thermal infrared emitters integrated into silicon photonic waveguides

Cost-efficient and easily integrable broadband mid-infrared (mid-IR) sources would significantly enhance the application space of photonic integrated circuits (PICs). Thermal incandescent sources are superior to other common mid-IR emitters based on semiconductor materials in terms of PIC compatibility, manufacturing costs, and bandwidth. Ideal thermal emitters would radiate directly into the desired modes of the PIC waveguides via near-field coupling and would be stable at very high temperatures. Graphene is a semi-metallic two-dimensional material with comparable emissivity to thin metallic thermal emitters. It allows maximum coupling into waveguides by placing it directly into their evanescent fields. Here, we demonstrate graphene mid-IR emitters integrated with photonic waveguides that couple directly into the fundamental mode of silicon waveguides designed for a wavelength of 4,2 {\mu}m relevant for CO${_2}$ sensing. High broadband emission intensity is observed at the waveguide-integrated graphene emitter. The emission at the output grating couplers confirms successful coupling into the waveguide mode. Thermal simulations predict emitter temperatures up to 1000{\deg}C, where the blackbody radiation covers the mid-IR region. A coupling efficiency {\eta}, defined as the light emitted into the waveguide divided by the total emission, of up to 68% is estimated, superior to data published for other waveguide-integrated emitters.

physics.optics

Analytics and Machine Learning Powered Wireless Network Optimization and Planning

It is important that the wireless network is well optimized and planned, using the limited wireless spectrum resources, to serve the explosively growing traffic and diverse applications needs of end users. Considering the challenges of dynamics and complexity of the wireless systems, and the scale of the networks, it is desirable to have solutions to automatically monitor, analyze, optimize, and plan the network. This article discusses approaches and solutions of data analytics and machine learning powered optimization and planning. The approaches include analyzing some important metrics of performances and experiences, at the lower layers and upper layers of open systems interconnection (OSI) model, as well as deriving a metric of the end user perceived network congestion indicator. The approaches include monitoring and diagnosis such as anomaly detection of the metrics, root cause analysis for poor performances and experiences. The approaches include enabling network optimization with tuning recommendations, directly targeting to optimize the end users experiences, via sensitivity modeling and analysis of the upper layer metrics of the end users experiences v.s. the improvement of the lower layers metrics due to tuning the hardware configurations. The approaches also include deriving predictive metrics for network planning, traffic demand distributions and trends, detection and prediction of the suppressed traffic demand, and the incentives of traffic gains if the network is upgraded. These approaches of optimization and planning are for accurate detection of optimization and upgrading opportunities at a large scale, enabling more effective optimization and planning such as tuning cells configurations, upgrading cells capacity with more advanced technologies or new hardware, adding more cells, etc., improving the network performances and providing better experiences to end users.

eess.SY

Three-dimensional printing of silica-glass structures with submicrometric features

Humanity's interest in manufacturing silica-glass objects extends back over three thousand years. Silica glass is resistant to heating and exposure to many chemicals, and it is transparent in a wide wavelength range. Due to these qualities, silica glass is used for a variety of applications that shape our modern life, such as optical fibers in medicine and telecommunications. However, its chemical stability and brittleness impede the structuring of silica glass, especially on the small scale. Techniques for three-dimensional (3D) printing of silica glass, such as stereolithography and direct ink writing, have recently been demonstrated, but the achievable minimum feature size is several tens of micrometers. While submicrometric silica-glass structures have many interesting applications, for example in micro-optics, they are currently manufactured using lithography techniques, which severely limits the 3D shapes that can be realized. Here, we show 3D printing of optically transparent silica-glass structures with submicrometric features. We achieve this by cross-linking hydrogen silsesquioxane to silica glass using nonlinear absorption of laser light followed by the dissolution of the unexposed material. We print a functional microtoroid resonator with out-of-plane fiber couplers to demonstrate the new possibilities for designing and building silica-glass microdevices in 3D.

physics.app-ph

Optimal User-Cell Association for 360 Video Streaming over Dense Wireless Networks

Delivering 360 degree video streaming for virtual and augmented reality presents many technical challenges especially in bandwidth starved wireless environments. Recently, a so-called two-tier approach has been proposed which delivers a basic-tier chunk and select enhancement-tier chunks to improve user experience while reducing network resources consumption. The video chunks are to be transmitted via unicast or multicast over an ultra-dense small cell infrastructure with enough bandwidth where small cells store video chunks in local caches. In this setup, user-cell association algorithms play a central role to efficiently deliver video since users may only download video chunks from the cell they are associated with. Motivated by this, we jointly formulate the problem of user-cell association and video chunk multicasting/unicasting as a mixed integer linear programming, prove its NP-hardness, and study the optimal solution via the Branch-and-Bound method. We then propose two polynomial-time, approximation algorithms and show via extensive simulations that they are near-optimal in practice and improve user experience by 30% compared to baseline user-cell association schemes.

cs.NI

DeepMVS: Learning Multi-view Stereopsis

We present DeepMVS, a deep convolutional neural network (ConvNet) for multi-view stereo reconstruction. Taking an arbitrary number of posed images as input, we first produce a set of plane-sweep volumes and use the proposed DeepMVS network to predict high-quality disparity maps. The key contributions that enable these results are (1) supervised pretraining on a photorealistic synthetic dataset, (2) an effective method for aggregating information across a set of unordered images, and (3) integrating multi-layer feature activations from the pre-trained VGG-19 network. We validate the efficacy of DeepMVS using the ETH3D Benchmark. Our results show that DeepMVS compares favorably against state-of-the-art conventional MVS algorithms and other ConvNet based methods, particularly for near-textureless regions and thin structures.

cs.CV