SearcharxivSearch

arXiv subjects

Lingfeng Wu

Publications and source records attributed to Lingfeng Wu.

3 recordsLinked to original sources

Heterogeneous back-end-of-line integration of thin-film lithium niobate on active silicon photonics for single-chip optical transceivers

The explosive growth of artificial intelligence, cloud computing, and large-scale machine learning is driving an urgent demand for short-reach optical interconnects featuring large bandwidth, low power consumption, high integration density, and low cost preferably adopting complementary metal-oxide-semiconductor (CMOS) processes. Heterogeneous integration of silicon photonics and thin-film lithium niobate (TFLN) combines the advantages of both platforms, and enables co-integration of high-performance modulators, photodetectors, and passive photonic components, offering an ideal route to meet these requirements. However, process incompatibilities have constrained the direct integration of TFLN with only passive silicon photonics. Here, we demonstrate the first heterogeneous back-end-of-line integration of TFLN with a full-functional and active silicon photonics platform via trench-based die-to-wafer bonding. This technology introduces TFLN after completing the full CMOS compatible processes for silicon photonics. Si/SiN passive components including low-loss fiber interfaces, 56-GHz Ge photodetectors, 100-GHz TFLN modulators, and multilayer metallization are integrated on a single silicon chip with efficient inter-layer and inter-material optical coupling. The integrated on-chip optical links exhibit greater than 60 GHz electrical-to-electrical bandwidth and support 128-GBaud OOK and 100-GBaud PAM4 transmission below forward error-correction thresholds, establishing a scalable platform for energy-efficient, high-capacity photonic systems.

physics.optics

Developing a Reliable, Fast, General-Purpose Hallucination Detection and Mitigation Service

Hallucination, a phenomenon where large language models (LLMs) produce output that is factually incorrect or unrelated to the input, is a major challenge for LLM applications that require accuracy and dependability. In this paper, we introduce a reliable and high-speed production system aimed at detecting and rectifying the hallucination issue within LLMs. Our system encompasses named entity recognition (NER), natural language inference (NLI), span-based detection (SBD), and an intricate decision tree-based process to reliably detect a wide range of hallucinations in LLM responses. Furthermore, we have crafted a rewriting mechanism that maintains an optimal mix of precision, response time, and cost-effectiveness. We detail the core elements of our framework and underscore the paramount challenges tied to response time, availability, and performance metrics, which are crucial for real-world deployment of these technologies. Our extensive evaluation, utilizing offline data and live production traffic, confirms the efficacy of our proposed framework and service.

cs.CL

Advances in Online Audio-Visual Meeting Transcription

This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved problem in realistic settings for over a decade. We show that this problem can be addressed by using a continuous speech separation approach. In addition, we describe an online audio-visual speaker diarization method that leverages face tracking and identification, sound source localization, speaker identification, and, if available, prior speaker information for robustness to various real world challenges. All components are integrated in a meeting transcription framework called SRD, which stands for "separate, recognize, and diarize". Experimental results using recordings of natural meetings involving up to 11 attendees are reported. The continuous speech separation improves a word error rate (WER) by 16.1% compared with a highly tuned beamformer. When a complete list of meeting attendees is available, the discrepancy between WER and speaker-attributed WER is only 1.0%, indicating accurate word-to-speaker association. This increases marginally to 1.6% when 50% of the attendees are unknown to the system.

eess.AS