SearcharxivSearch

arXiv subjects

Hanyu Zheng

Publications and source records attributed to Hanyu Zheng.

11 recordsLinked to original sources

Improving Large Vision-Language Models' Understanding for Flow Field Data

Large Vision-Language Models (LVLMs) have shown impressive capabilities across a range of tasks that integrate visual and textual understanding, such as image captioning and visual question answering. These models are trained on large-scale image and video datasets paired with text, enabling them to bridge visual perception and natural language processing. However, their application to scientific domains, especially in interpreting complex field data commonly used in the natural sciences, remains underexplored. In this work, we introduce FieldLVLM, a novel framework designed to improve large vision-language models' understanding of field data. FieldLVLM consists of two main components: a field-aware language generation strategy and a data-compressed multimodal model tuning. The field-aware language generation strategy leverages a special-purpose machine learning pipeline to extract key physical features from field data, such as flow classification, Reynolds number, and vortex patterns. This information is then converted into structured textual descriptions that serve as a dataset. The data-compressed multimodal model tuning focuses on LVLMs with these generated datasets, using a data compression strategy to reduce the complexity of field inputs and retain only the most informative values. This ensures compatibility with the models language decoder and guides its learning more effectively. Experimental results on newly proposed benchmark datasets demonstrate that FieldLVLM significantly outperforms existing methods in tasks involving scientific field data. Our findings suggest that this approach opens up new possibilities for applying large vision-language models to scientific research, helping bridge the gap between large models and domain-specific discovery.

cs.CV

Ultrawide-angle diffraction-limited 2D beam steering via hybrid integrated metasurface-photonic circuit

Two-dimensional (2D) wide field-of-view (FOV) beam steering is a key enabling capability for emerging free-space optical systems, including inter-satellite optical links, airborne LiDAR, point-to-point optical wireless communications, and collaborative robotic platforms. These applications require rapid acquisition and tracking across both azimuth and elevation; architectures that offer wide scanning in only one dimension while maintaining limited coverage in the orthogonal direction constrain link availability, coverage uniformity, and system agility. Here, we demonstrate a chip-scale platform for ultrawide-angle, diffraction-limited 2D beam steering based on hybrid integration of a silicon photonic integrated circuit (PIC) and an optical metasurface. A free-form micro-optical reflector efficiently transforms the guided waveguide mode into an expanded free-space beam that illuminates an analytically optimized ultrawide-FOV metasurface. The integrated system achieves a measured FOV exceeding 160° while maintaining diffraction-limited beam quality over a broad angular range at telecom wavelengths. This hybrid PIC-metasurface architecture provides a compact and scalable route to high-quality 2D beam steering and establishes a practical pathway toward integrated optical projectors for space-based optical communications and other applications requiring agile, wide-angle, high-fidelity beam control.

physics.optics

Wafer-scale conformal metasurface optics

Curved and conformal optics offer significant advantages by unlocking additional geometric degrees of freedom for optical design. These capabilities enable enhanced optical performance and are essential for meeting non-optical constraints, such as those imposed by ergonomics, aerodynamics, or wearability. However, existing fabrication techniques such as direct electron or laser beam writing on curved substrates, and soft-stamp-based transfer or nanoimprint lithography suffer from limitations in scalability, yield, geometry control, and alignment accuracy. Here, we present a scalable fabrication strategy for curved and conformal metasurface optics leveraging thermoforming, an industry-standard, high-throughput manufacturing process widely used for shaping thermoplastics. Our approach uniquely enables wafer-scale production of highly curved metasurface optics, achieving sub-millimeter radii of curvature and micron-level alignment precision. To guide the design and fabrication process, we developed a thermorheological model that accurately predicts and compensates for the large strains induced during thermoforming. This allows for precise control of metasurface geometry and preservation of optical function, yielding devices with diffraction-limited performance. As a demonstration, we implemented an artificial compound eye comprising freeform micro-metalens arrays. Compared to traditional micro-optical counterparts, the device exhibits an expanded field of view, reduced aberrations, and improved uniformity, highlighting the potential of thermoformed metasurfaces for next-generation optical systems.

physics.optics

Digital Modeling on Large Kernel Metamaterial Neural Network

Deep neural networks (DNNs) utilized recently are physically deployed with computational units (e.g., CPUs and GPUs). Such a design might lead to a heavy computational burden, significant latency, and intensive power consumption, which are critical limitations in applications such as the Internet of Things (IoT), edge computing, and the usage of drones. Recent advances in optical computational units (e.g., metamaterial) have shed light on energy-free and light-speed neural networks. However, the digital design of the metamaterial neural network (MNN) is fundamentally limited by its physical limitations, such as precision, noise, and bandwidth during fabrication. Moreover, the unique advantages of MNN's (e.g., light-speed computation) are not fully explored via standard 3x3 convolution kernels. In this paper, we propose a novel large kernel metamaterial neural network (LMNN) that maximizes the digital capacity of the state-of-the-art (SOTA) MNN with model re-parametrization and network compression, while also considering the optical limitation explicitly. The new digital learning scheme can maximize the learning capacity of MNN while modeling the physical restrictions of meta-optic. With the proposed LMNN, the computation cost of the convolutional front-end can be offloaded into fabricated optical hardware. The experimental results on two publicly available datasets demonstrate that the optimized hybrid design improved classification accuracy while reducing computational latency. The development of the proposed LMNN is a promising step towards the ultimate goal of energy-free and light-speed AI.

cs.CV

Intelligent Multi-channel Meta-imagers for Accelerating Machine Vision

Rapid developments in machine vision have led to advances in a variety of industries, from medical image analysis to autonomous systems. These achievements, however, typically necessitate digital neural networks with heavy computational requirements, which are limited by high energy consumption and further hinder real-time decision-making when computation resources are not accessible. Here, we demonstrate an intelligent meta-imager that is designed to work in concert with a digital back-end to off-load computationally expensive convolution operations into high-speed and low-power optics. In this architecture, metasurfaces enable both angle and polarization multiplexing to create multiple information channels that perform positive and negatively valued convolution operations in a single shot. The meta-imager is employed for object classification, experimentally achieving 98.6% accurate classification of handwritten digits and 88.8% accuracy in classifying fashion images. With compactness, high speed, and low power consumption, this approach could find a wide range of applications in artificial intelligence and machine vision applications.

cs.CV

Enhanced compound-protein binding affinity prediction by representing protein multimodal information via a coevolutionary strategy

Due to the lack of a method to efficiently represent the multimodal information of a protein, including its structure and sequence information, predicting compound-protein binding affinity (CPA) still suffers from low accuracy when applying machine learning methods. To overcome this limitation, in a novel end-to-end architecture (named FeatNN), we develop a coevolutionary strategy to jointly represent the structure and sequence features of proteins and ultimately optimize the mathematical models for predicting CPA. Furthermore, from the perspective of data-driven approach, we proposed a rational method that can utilize both high- and low-quality databases to optimize the accuracy and generalization ability of FeatNN in CPA prediction tasks. Notably, we visually interpret the feature interaction process between sequence and structure in the rationally designed architecture. As a result, FeatNN considerably outperforms the state-of-the-art (SOTA) baseline in virtual drug screening tasks, indicating the feasibility of this approach for practical use. FeatNN provides an outstanding method for higher CPA prediction accuracy and better generalization ability by efficiently representing multimodal information of proteins via a coevolutionary strategy.

q-bio.BM

Meta-optic Accelerators for Object Classifiers

Rapid advances in deep learning have led to paradigm shifts in a number of fields, from medical image analysis to autonomous systems. These advances, however, have resulted in digital neural networks with large computational requirements, resulting in high energy consumption and limitations in real-time decision making when computation resources are limited. Here, we demonstrate a meta-optic based neural network accelerator that can off-load computationally expensive convolution operations into high-speed and low-power optics. In this architecture, metasurfaces enable both spatial multiplexing and additional information channels, such as polarization, in object classification. End-to-end design is used to co-optimize the optical and digital systems resulting in a robust classifier that achieves 95% accurate classification of handwriting digits and 94% accuracy in classifying both the digit and its polarization state. This approach could enable compact, high-speed, and low-power image and information processing systems for a wide range of applications in machine-vision and artificial intelligence.

physics.optics

All-Dielectric Meta-optics for High-Efficiency Independent Amplitude and Phase Manipulation

Metasurfaces, composed of subwavelength scattering elements, have demonstrated remarkable control over the transmitted amplitude, phase, and polarization of light. However, manipulating the amplitude upon transmission has required loss if a single metasurface is used. Here, we describe high-efficiency independent manipulation of the amplitude and phase of a beam using two lossless phase-only metasurfaces separated by a distance. With this configuration, we experimentally demonstrate optical components such as combined beam-forming and splitting devices, as well as those for forming complex-valued, three-dimensional holograms. The compound meta-optic platform provides a promising approach for achieving high performance optical holographic displays and compact optical components, while exhibiting a high overall efficiency.

physics.optics

Image Processing Based on Compound Flat Optics

Image processing has become a critical technology in a variety of science and engineering disciplines. While most image processing is performed digitally, optical analog processing has the advantages of being low-power and high-speed though it requires a large volume. Here, we demonstrate optical analog imaging processing using a flat optic for direct image differentiation allowing one to significantly shrink the required optical system size. We first demonstrate how the image differentiator can be combined with traditional imaging systems such as a commercial optical microscope and camera sensor for edge detection. Second, we demonstrate how the entire analog processing system can be realized as a monolithic compound flat optic by integrating the differentiator with a metalens. The compound nanophotonic system manifests the advantage of thin form factor optics as well as the ability to implement complex transfer functions and could open new opportunities in applications such as biological imaging and machine vision.

physics.optics

Ultra-thin, High-efficiency Mid-Infrared Transmissive Huygens Meta-Optics

The mid-infrared (mid-IR) is a strategically important band for numerous applications ranging from night vision to biochemical sensing. Unlike visible or near-infrared optical parts which are commonplace and economically available off-the-shelf, mid-IR optics often requires exotic materials or complicated processing, which accounts for their high cost and inferior quality compared to their visible or near-infrared counterparts. Here we theoretically analyzed and experimentally realized a Huygens metasurface platform capable of fulfilling a diverse cross-section of optical functions in the mid-IR. The meta-optical elements were constructed using high-index chalcogenide films deposited on fluoride substrates:the choices of wide-band transparent materials allow the design to be scaled across a broad infrared spectrum. Capitalizing on a novel two-component Huygens' meta-atom design, the meta-optical devices feature an ultra-thin profile ($λ_0/8$ in thickness, where $λ_0$ is the free-space wavelength) and measured optical efficiencies up to 75% in transmissive mode, both of which represent major improvements over state-of-the-art. We have also demonstrated, for the first time, mid-IR transmissive meta-lenses with diffraction-limited focusing and imaging performance. The projected size, weight and power advantages, coupled with the manufacturing scalability leveraging standard microfabrication technologies, make the Huygens meta-optical devices promising for next-generation mid-IR system applications.

physics.optics

Chalcogenide Glass-on-Graphene Photonics

Two-dimensional (2-D) materials are of tremendous interest to integrated photonics given their singular optical characteristics spanning light emission, modulation, saturable absorption, and nonlinear optics. To harness their optical properties, these atomically thin materials are usually attached onto prefabricated devices via a transfer process. In this paper, we present a new route for 2-D material integration with planar photonics. Central to this approach is the use of chalcogenide glass, a multifunctional material which can be directly deposited and patterned on a wide variety of 2-D materials and can simultaneously function as the light guiding medium, a gate dielectric, and a passivation layer for 2-D materials. Besides claiming improved fabrication yield and throughput compared to the traditional transfer process, our technique also enables unconventional multilayer device geometries optimally designed for enhancing light-matter interactions in the 2-D layers. Capitalizing on this facile integration method, we demonstrate a series of high-performance glass-on-graphene devices including ultra-broadband on-chip polarizers, energy-efficient thermo-optic switches, as well as graphene-based mid-infrared (mid-IR) waveguide-integrated photodetectors and modulators.

physics.optics