SearcharxivSearch

arXiv subjects

Ethan Tseng

Publications and source records attributed to Ethan Tseng.

11 recordsLinked to original sources

Intrinsic Limitations of Single Layer Polychromatic Metalens for Virtual Reality Visors

Virtual and augmented reality (VR/AR) visors require compact and lightweight optics. Metalenses have been widely proposed as ultrathin replacements for bulky refractive eyepieces, with performance typically assessed using point spread function (PSF) and modulation transfer function (MTF) measurements. Here, we design, fabricate, and characterize a single-layer silicon nitride metalens optimized for the three emission peaks of an RGB OLED display, and benchmark it against refractive and Fresnel eyepieces. Under coherent illumination, the metalens exhibits a tightly confined PSF and strong mid-to-high spatial-frequency MTF, suggesting excellent optical performance. However, when evaluated in a realistic system-level VR testbed incorporating incoherent OLED illumination, a dynamic-pupil eye model, and near-eye-relevant focal lengths, the same device exhibits pronounced ghosting and background haze. We show that these artifacts arise from the intrinsic multifocal nature of polychromatic diffractive focusing and demonstrate that common mitigation strategies such as narrowband filtering and long-focal-length relay optics merely mask, rather than resolve, the issue. Our results establish that meta-optics for AR/VR must be evaluated under realistic system-level conditions to reveal their true imaging performance.

physics.optics

Lucky High Dynamic Range Smartphone Imaging

While the human eye can perceive an impressive twenty stops of dynamic range, smartphone camera sensors remain limited to about twelve stops despite decades of research. A variety of high dynamic range (HDR) image capture and processing techniques have been proposed, and, in practice, they can extend the dynamic range by 3-5 stops for handheld photography. This paper proposes an approach that robustly captures dynamic range using a handheld smartphone camera and lightweight networks suitable for running on mobile devices. Our method operates indirectly on linear raw pixels in bracketed exposures. Every pixel in the final HDR image is a convex combination of input pixels in the neighborhood, adjusted for exposure, and thus avoids hallucination artifacts typical of recent deep image synthesis networks. We validate our system on both synthetic imagery and unseen real bracketed images -- we confirm zero-shot generalization of the method to smartphone camera captures. Our iterative inference architecture is capable of processing an arbitrary number of bracketed input photos, and we show examples from capture stacks containing 3--9 images. Our training process relies only on synthetic captures yet generalizes to unseen real photos from several cameras. Moreover, we show that this training scheme improves other SOTA methods over their pretrained counterparts.

cs.CV

Neural Étendue Expander for Ultra-Wide-Angle High-Fidelity Holographic Display

Holographic displays can generate light fields by dynamically modulating the wavefront of a coherent beam of light using a spatial light modulator, promising rich virtual and augmented reality applications. However, the limited spatial resolution of existing dynamic spatial light modulators imposes a tight bound on the diffraction angle. As a result, modern holographic displays possess low étendue, which is the product of the display area and the maximum solid angle of diffracted light. The low étendue forces a sacrifice of either the field-of-view (FOV) or the display size. In this work, we lift this limitation by presenting neural étendue expanders. This new breed of optical elements, which is learned from a natural image dataset, enables higher diffraction angles for ultra-wide FOV while maintaining both a compact form factor and the fidelity of displayed contents to human viewers. With neural étendue expanders, we experimentally achieve 64$\times$ étendue expansion of natural images in full color, expanding the FOV by an order of magnitude horizontally and vertically, with high-fidelity reconstruction quality (measured in PSNR) over 29 dB on retinal-resolution images.

eess.IV

Beating bandwidth limits for large aperture broadband nano-optics

Flat optics have been proposed as an attractive approach for the implementation of new imaging and sensing modalities to replace and augment refractive optics. However, chromatic aberrations impose fundamental limitations on diffractive flat optics. As such, true broadband high-quality imaging has thus far been out of reach for low f-number, large aperture, flat optics. In this work, we overcome these intrinsic fundamental limitations, achieving broadband imaging in the visible wavelength range with a flat meta-optic, co-designed with computational reconstruction. We derive the necessary conditions for a broadband, 1 cm aperture, f/2 flat optic, with a diagonal field of view of 30° and an average system MTF contrast of 30% or larger for a spatial frequency of 100 lp/mm in the visible band (> 50 % for 70 lp/mm and below). Finally, we use a coaxial, dual-aperture system to train the broadband imaging meta-optic with a learned reconstruction method operating on pair-wise captured imaging data. Fundamentally, our work challenges the entrenched belief of the inability of capturing high-quality, full-color images using a single large aperture meta-optic.

physics.optics

Spatially Varying Nanophotonic Neural Networks

The explosive growth of computation and energy cost of artificial intelligence has spurred strong interests in new computing modalities as potential alternatives to conventional electronic processors. Photonic processors that execute operations using photons instead of electrons, have promised to enable optical neural networks with ultra-low latency and power consumption. However, existing optical neural networks, limited by the underlying network designs, have achieved image recognition accuracy far below that of state-of-the-art electronic neural networks. In this work, we close this gap by embedding massively parallelized optical computation into flat camera optics that perform neural network computation during the capture, before recording an image on the sensor. Specifically, we harness large kernels and propose a large-kernel spatially-varying convolutional neural network learned via low-dimensional reparameterization techniques. We experimentally instantiate the network with a flat meta-optical system that encompasses an array of nanophotonic structures designed to induce angle-dependent responses. Combined with an extremely lightweight electronic backend with approximately 2K parameters we demonstrate a reconfigurable nanophotonic neural network reaches 72.76\% blind test classification accuracy on CIFAR-10 dataset, and, as such, the first time, an optical neural network outperforms the first modern digital neural network -- AlexNet (72.64\%) with 57M parameters, bringing optical neural network into modern deep learning era.

cs.CV

Stochastic Light Field Holography

The Visual Turing Test is the ultimate goal to evaluate the realism of holographic displays. Previous studies have focused on addressing challenges such as limited étendue and image quality over a large focal volume, but they have not investigated the effect of pupil sampling on the viewing experience in full 3D holograms. In this work, we tackle this problem with a novel hologram generation algorithm motivated by matching the projection operators of incoherent Light Field and coherent Wigner Function light transport. To this end, we supervise hologram computation using synthesized photographs, which are rendered on-the-fly using Light Field refocusing from stochastically sampled pupil states during optimization. The proposed method produces holograms with correct parallax and focus cues, which are important for passing the Visual Turing Test. We validate that our approach compares favorably to state-of-the-art CGH algorithms that use Light Field and Focal Stack supervision. Our experiments demonstrate that our algorithm significantly improves the realism of the viewing experience for a variety of different pupil states.

cs.CV

Pupil-aware Holography

Holographic displays promise to deliver unprecedented display capabilities in augmented reality applications, featuring a wide field of view, wide color gamut, spatial resolution, and depth cues all in a compact form factor. While emerging holographic display approaches have been successful in achieving large etendue and high image quality as seen by a camera, the large etendue also reveals a problem that makes existing displays impractical: the sampling of the holographic field by the eye pupil. Existing methods have not investigated this issue due to the lack of displays with large enough etendue, and, as such, they suffer from severe artifacts with varying eye pupil size and location. We show that the holographic field as sampled by the eye pupil is highly varying for existing display setups, and we propose pupil-aware holography that maximizes the perceptual image quality irrespective of the size, location, and orientation of the eye pupil in a near-eye holographic display. We validate the proposed approach both in simulations and on a prototype holographic display and show that our method eliminates severe artifacts and significantly outperforms existing approaches.

cs.GR

ZeroScatter: Domain Transfer for Long Distance Imaging and Vision through Scattering Media

Adverse weather conditions, including snow, rain, and fog, pose a major challenge for both human and computer vision. Handling these environmental conditions is essential for safe decision making, especially in autonomous vehicles, robotics, and drones. Most of today's supervised imaging and vision approaches, however, rely on training data collected in the real world that is biased towards good weather conditions, with dense fog, snow, and heavy rain as outliers in these datasets. Without training data, let alone paired data, existing autonomous vehicles often limit themselves to good conditions and stop when dense fog or snow is detected. In this work, we tackle the lack of supervised training data by combining synthetic and indirect supervision. We present ZeroScatter, a domain transfer method for converting RGB-only captures taken in adverse weather into clear daytime scenes. ZeroScatter exploits model-based, temporal, multi-view, multi-modal, and adversarial cues in a joint fashion, allowing us to train on unpaired, biased data. We assess the proposed method on in-the-wild captures, and the proposed method outperforms existing monocular descattering approaches by 2.8 dB PSNR on controlled fog chamber measurements.

cs.CV

Neural Nano-Optics for High-quality Thin Lens Imaging

Nano-optic imagers that modulate light at sub-wavelength scales could unlock unprecedented applications in diverse domains ranging from robotics to medicine. Although metasurface optics offer a path to such ultra-small imagers, existing methods have achieved image quality far worse than bulky refractive alternatives, fundamentally limited by aberrations at large apertures and low f-numbers. In this work, we close this performance gap by presenting the first neural nano-optics. We devise a fully differentiable learning method that learns a metasurface physical structure in conjunction with a novel, neural feature-based image reconstruction algorithm. Experimentally validating the proposed method, we achieve an order of magnitude lower reconstruction error. As such, we present the first high-quality, nano-optic imager that combines the widest field of view for full-color metasurface operation while simultaneously achieving the largest demonstrated 0.5 mm, f/2 aperture.

physics.optics

Gated3D: Monocular 3D Object Detection From Temporal Illumination Cues

Today's state-of-the-art methods for 3D object detection are based on lidar, stereo, or monocular cameras. Lidar-based methods achieve the best accuracy, but have a large footprint, high cost, and mechanically-limited angular sampling rates, resulting in low spatial resolution at long ranges. Recent approaches based on low-cost monocular or stereo cameras promise to overcome these limitations but struggle in low-light or low-contrast regions as they rely on passive CMOS sensors. In this work, we propose a novel 3D object detection modality that exploits temporal illumination cues from a low-cost monocular gated imager. We propose a novel deep detector architecture, Gated3D, that is tailored to temporal illumination cues from three gated images. Gated images allow us to exploit mature 2D object feature extractors that guide the 3D predictions through a frustum segment estimation. We assess the proposed method on a novel 3D detection dataset that includes gated imagery captured in over 10,000 km of driving data. We validate that our method outperforms state-of-the-art monocular and stereo approaches at long distances. We will release our code and dataset, opening up a new sensor modality as an avenue to replace lidar in autonomous driving.

cs.CV

Automatic In-line Quantitative Myocardial Perfusion Mapping: processing algorithm and implementation

Quantitative myocardial perfusion mapping has advantages over qualitative assessment, including the ability to detect global flow reduction. However, it is not clinically available and remains as a research tool. Building upon the previously described imaging sequence, this paper presents algorithm and implementation of an automated solution for inline perfusion flow mapping with step by step performance characterization. An inline perfusion flow mapping workflow is proposed and demonstrated on normal volunteers. Initial evaluation demonstrates the fully automated proposed solution for the respiratory motion correction, AIF LV mask detection and pixel-wise mapping, from free-breathing myocardial perfusion imaging.

eess.IV