SearcharxivSearch

arXiv subjects

Junwen He

Publications and source records attributed to Junwen He.

5 recordsLinked to original sources

Low-Loss Integration of High-Density Polymer Waveguides with Silicon Photonics for Co-Packaged Optics

Co-Packaged Optics applications require scalable and high-yield optical interfacing solutions to silicon photonic chiplets, offering low-loss, broadband, and polarization-independent optical coupling while maintaining compatibility with widely used approaches for electrical redistribution. We present two heterogeneous integration techniques that enable high-density electrical and optical I/O connections, utilizing adiabatic coupling between on-chip silicon nitride (SiN) waveguides and package-level polymer optical waveguides. In the first approach, polymer waveguides are patterned using standard lithography directly on the surface of the photonic chip, ensuring compatibility with chip embedding as commonly employed in chip-first fanout wafer-level packaging. In the second approach, photonic chips are flip-chip bonded to the package substrate. Both techniques have been experimentally validated, achieving a coupling efficiency near 1 dB between SiN and polymer waveguides in O-band, for both TE and TM polarizations. SiN tapers were designed using the "Mono" method to optimize phase-matching conditions between the two waveguides, a critical requirement for integrating diverse optical components. These results demonstrate the potential of polymer waveguides in Co-Packaged Optics applications, achieving sub-2 dB chip-to-chip and chip-to-fiber coupling losses.

physics.optics

GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts

Text logo design heavily relies on the creativity and expertise of professional designers, in which arranging element layouts is one of the most important procedures. However, this specific task has received limited attention, often overshadowed by broader layout generation tasks such as document or poster design. In this paper, we propose a Vision-Language Model (VLM)-based framework that generates content-aware text logo layouts by integrating multi-modal inputs with user-defined constraints, enabling more flexible and robust layout generation for real-world applications. We introduce two model techniques that reduce the computational cost for processing multiple glyph images simultaneously, without compromising performance. To support instruction tuning of our model, we construct two extensive text logo datasets that are five times larger than existing public datasets. In addition to geometric annotations (\textit{e.g.}, text masks and character recognition), our datasets include detailed layout descriptions in natural language, enabling the model to reason more effectively in handling complex designs and custom user inputs. Experimental results demonstrate the effectiveness of our proposed framework and datasets, outperforming existing methods on various benchmarks that assess geometric aesthetics and human preferences.

cs.CV

Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception

Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However, there still remains a gap in providing fine-grained pixel-level perceptions and extending interactions beyond text-specific inputs. In this work, we propose {\bf{AnyRef}}, a general MLLM model that can generate pixel-wise object perceptions and natural language descriptions from multi-modality references, such as texts, boxes, images, or audio. This innovation empowers users with greater flexibility to engage with the model beyond textual and regional prompts, without modality-specific designs. Through our proposed refocusing mechanism, the generated grounding output is guided to better focus on the referenced object, implicitly incorporating additional pixel-level supervision. This simple modification utilizes attention scores generated during the inference of LLM, eliminating the need for extra computations while exhibiting performance enhancements in both grounding masks and referring expressions. With only publicly available training data, our model achieves state-of-the-art results across multiple benchmarks, including diverse modality referring segmentation and region-level referring expression generation.

cs.CV

Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning

Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits their potential for exploring cross-domain information. We propose a deeply unified framework for depth-aware panoptic segmentation, which performs joint segmentation and depth estimation both in a per-segment manner with identical object queries. To narrow the gap between the two tasks, we further design a geometric query enhancement method, which is able to integrate scene geometry into object queries using latent representations. In addition, we propose a bi-directional guidance learning approach to facilitate cross-task feature learning by taking advantage of their mutual relations. Our method sets the new state of the art for depth-aware panoptic segmentation on both Cityscapes-DVPS and SemKITTI-DVPS datasets. Moreover, our guidance learning approach is shown to deliver performance improvement even under incomplete supervision labels.

cs.CV

Micro-optical Tandem Luminescent Solar Concentrators

Traditional concentrating photovoltaic (CPV) systems utilize multijunction cells to minimize thermalization losses, but cannot efficiently capture diffuse sunlight, which contributes to a high levelized cost of energy (LCOE) and limits their use to geographical regions with high direct sunlight insolation. Luminescent solar concentrators (LSCs) harness light generated by luminophores embedded in a light-trapping waveguide to concentrate light onto smaller cells. LSCs can absorb both direct and diffuse sunlight, and thus can operate as flat plate receivers at a fixed tilt and with a conventional module form factor. However, current LSCs experience significant power loss through parasitic luminophore absorption and incomplete light trapping by the optical waveguide. Here we introduce a tandem LSC device architecture that overcomes both of these limitations, consisting of a PLMA polymer layer with embedded CdSe/CdS quantum dot (QD) luminophores and InGaP micro-cells, which serve as a high bandgap absorber on top of a conventional Si photovoltaic. We experimentally synthesize CdSe/CdS QDs with exceptionally high quantum-yield (99%) and ultra-narrowband emission optimally matched to fabricated III-V InGaP micro-cells. Using a Monte Carlo ray-tracing model, we show the radiative limit power conversion efficiency for a module with these components to be 30.8% diffuse sunlight conditions. These results indicate that a tandem LSC-on-Si architecture could significantly improve upon the efficiency of a conventional Si photovoltaic module with simple and straightforward alterations of the module lamination steps of a Si photovoltaic manufacturing process, with promise for widespread module deployment across diverse geographical regions and energy markets.

physics.app-ph