SearcharxivSearch

arXiv subjects

Hongkuan Zhang

Publications and source records attributed to Hongkuan Zhang.

7 recordsLinked to original sources

Tailoring reflectionless complex media for non-Abelian braiding of acoustic modes

Multiple scattering of sound and light can be tailored for diverse applications. Despite the great progress enabled by technologies such as time-reversal propagation and wavefront shaping, the full control of the transmission matrix remains a significant challenge. In this work, we propose a multi-scattering-based approach to design reflectionless complex media with an arbitrary unitary transmission matrix. As such, the perfect transmission of waves through such a medium performs a unitary operation. Based on this principle, we experimentally demonstrated braiding of multiple waveguide modes in an acoustic waveguide via multiple scattering and showed non-Abelian characteristics arising from the concatenation of distinct complex media. Furthermore, we show that the principle can be extended for realizing arbitrary unitary operations beyond braiding. Our scheme uses generalized Wigner-Smith operators to design the optimal acoustic complex media with near-arbitrary targeted functionalities. The scheme is generally applicable beyond acoustics, with broad implications to other wave types. Our results demonstrate unprecedented control over multiple-scattering waves and establish complex media as a viable route to modal control and operations on compact platforms, providing a new design paradigm for multimode, reconfigurable wave-based devices with potential applications in multiplexed communication, imaging, and the manipulation of quantum waves.

physics.app-ph

Refine Medical Diagnosis Using Generation Augmented Retrieval and Clinical Practice Guidelines

Current medical language models, adapted from large language models (LLMs), typically predict ICD code-based diagnosis from electronic health records (EHRs) because these labels are readily available. However, ICD codes do not capture the nuanced, context-rich reasoning clinicians use for diagnosis. Clinicians synthesize diverse patient data and reference clinical practice guidelines (CPGs) to make evidence-based decisions. This misalignment limits the clinical utility of existing models. We introduce GARMLE-G, a Generation-Augmented Retrieval framework that grounds medical language model outputs in authoritative CPGs. Unlike conventional Retrieval-Augmented Generation based approaches, GARMLE-G enables hallucination-free outputs by directly retrieving authoritative guideline content without relying on model-generated text. It (1) integrates LLM predictions with EHR data to create semantically rich queries, (2) retrieves relevant CPG knowledge snippets via embedding similarity, and (3) fuses guideline content with model output to generate clinically aligned recommendations. A prototype system for hypertension diagnosis was developed and evaluated on multiple metrics, demonstrating superior retrieval precision, semantic relevance, and clinical guideline adherence compared to RAG-based baselines, while maintaining a lightweight architecture suitable for localized healthcare deployment. This work provides a scalable, low-cost, and hallucination-free method for grounding medical language models in evidence-based clinical practice, with strong potential for broader clinical deployment.

cs.CL

Optimizing multi-user indoor sound communications with acoustic reconfigurable metasurfaces

Sound in indoor spaces forms a complex wavefield due to multiple scattering encountered by the sound. Indoor acoustic communication involving multiple sources and receivers thus inevitably suffers from cross-talks. Here, we demonstrate the isolation of acoustic communication channels in a room by wavefield shaping using acoustic reconfigurable metasurfaces (ARMs) controlled by optimization protocols based on communication theories. The ARMs have 200 electrically switchable units, each selectively offering 0 or π phase shifts in the reflected waves. The sound field is reshaped for maximal Shannon capacity and minimal cross-talk simultaneously. We demonstrate diverse acoustic functionalities over a spectrum much larger than the coherence bandwidth of the room, including multi-channel, multi-spectral channel isolations, and frequency-multiplexed acoustic communication. Our work shows that wavefield shaping in complex media can offer new strategies for future acoustic engineering.

cs.SD

Extending TrOCR for Text Localization-Free OCR of Full-Page Scanned Receipt Images

Digitization of scanned receipts aims to extract text from receipt images and save it into structured documents. This is usually split into two sub-tasks: text localization and optical character recognition (OCR). Most existing OCR models only focus on the cropped text instance images, which require the bounding box information provided by a text region detection model. Introducing an additional detector to identify the text instance images in advance adds complexity, however instance-level OCR models have very low accuracy when processing the whole image for the document-level OCR, such as receipt images containing multiple text lines arranged in various layouts. To this end, we propose a localization-free document-level OCR model for transcribing all the characters in a receipt image into an ordered sequence end-to-end. Specifically, we finetune the pretrained instance-level model TrOCR with randomly cropped image chunks, and gradually increase the image chunk size to generalize the recognition ability from instance images to full-page images. In our experiments on the SROIE receipt OCR dataset, the model finetuned with our strategy achieved 64.4 F1-score and a 22.8% character error rate (CER), respectively, which outperforms the baseline results with 48.5 F1-score and 50.6% CER. The best model, which splits the full image into 15 equally sized chunks, gives 87.8 F1-score and 4.98% CER with minimal additional pre or post-processing of the output. Moreover, the characters in the generated document-level sequences are arranged in the reading order, which is practical for real-world applications.

cs.CL

Cross-Modal Similarity-Based Curriculum Learning for Image Captioning

Image captioning models require the high-level generalization ability to describe the contents of various images in words. Most existing approaches treat the image-caption pairs equally in their training without considering the differences in their learning difficulties. Several image captioning approaches introduce curriculum learning methods that present training data with increasing levels of difficulty. However, their difficulty measurements are either based on domain-specific features or prior model training. In this paper, we propose a simple yet efficient difficulty measurement for image captioning using cross-modal similarity calculated by a pretrained vision-language model. Experiments on the COCO and Flickr30k datasets show that our proposed approach achieves superior performance and competitive convergence speed to baselines without requiring heuristics or incurring additional training costs. Moreover, the higher model performance on difficult examples and unseen data also demonstrates the generalization ability.

cs.CV

Physical rendering of synthetic spaces for topological sound transport

Synthetic dimensions can be rendered in the physical space and this has been achieved with photonics and cold atomic gases, however, little to no work has been succeeded in acoustics because acoustic wave-guides cannot be weakly coupled in a continuous fashion. Here, we establish the theoretical principles and for the first time manufacture acoustic crystals composed of arrays of acoustic cavities strongly coupled through modulated channels to evidence one-dimensional (1D) and two-dimensional (2D) dynamic topological pumpings. In particular, the topological edge-bulkedge and corner-bulk-corner transport are physically illustrated in finite-sized acoustic structures. We delineate the generated 2D and four-dimensional (4D) quantum Hall effects by calculating first and second Chern numbers and demonstrating robustness against the geometrical imperfections. Synthetic dimensions could provide a powerful way for acoustic topological wave steering and open up a new platform to explore higher-order topological matter in dimensions four and higher.

cond-mat.mes-hall

An asymmetric elastic metamaterial model for elastic wave cloaking

Elastic material with its elastic tensor losing minor symmetry is considered impossible without introducing artificially body torque. Here we demonstrate the feasibility of such material by introducing rotational resonance, the amplified rotational inertia of the microstructure during dynamical loading breaks naturally the shear stress symmetry, without resorting to external body torque or any other active means. This concept is illustrated through a realistic mass-spring model together with analytical homogenization technique and band structure analysis. It is also proven that this metamaterial model can be deliberately tuned to meet the material requirement defined by transformation method for full control of elastic wave, and the relation bridging the microstructure and the desired wave functionality is explicitly given. Application of this asymmetric metamaterial to design elastic wave cloak is demonstrated and validated by numerical simulation. The study paves the way for material design used to construct the transformation media for controlling elastic wave and related devices.

physics.class-ph