SearcharxivSearch

arXiv subjects

Akira Kawai

Publications and source records attributed to Akira Kawai.

10 recordsLinked to original sources

Ghost infrared spectroscopy with bright twin beams

Frequency-correlated light offers a route to mid-infrared (MIR) spectroscopy without direct spectral detection in the MIR. Previous MIR ghost spectroscopy has mainly relied on low-gain spontaneous parametric down-conversion (SPDC) and photon-pair coincidence measurements, where the limited photon flux has restricted acquisition times to longer than one minute. Here, we demonstrate ghost infrared spectroscopy using bright twin beams generated by high-gain parametric down-conversion (PDC). High-gain PDC amplifies vacuum fluctuations, producing a different pair of frequency-correlated random spectra in each pump pulse. This pulse-resolved stochastic emission is naturally matched to time-stretch detection, which records the spectrum of the correlated near-infrared telecom signal for every pulse, while the MIR idler is measured by bucket detection. Consequently, each pump pulse yields one paired projection measurement comprising a spectrally resolved reference and the corresponding bucket value. We reconstruct the transmission spectrum of a structured optical filter and the molecular vibrational absorption spectrum of liquid benzene near 3.3 um with millisecond-scale acquisition, in good agreement with Fourier-transform infrared spectroscopy. This reduces the acquisition time by four to five orders of magnitude compared with previous MIR ghost spectroscopy demonstrations. These results transform ghost spectroscopy from coincidence-based photon counting to high-flux analog correlation spectroscopy, establishing a practical architecture for high-speed computational infrared spectroscopy driven by a narrowband semiconductor laser.

physics.optics

LLMs Capture Emotion Labels, Not Emotion Uncertainty: Distributional Analysis and Calibration of Human-LLM Judgment Gaps

Human annotators frequently disagree on emotion labels, yet most evaluations of Large Language Model (LLM) emotion annotation collapse these judgments into a single gold standard, discarding the distributional information that disagreement encodes. We ask whether LLMs capture the structure of this disagreement, not just majority labels, by comparing emotion judgment distributions between human annotators and four zero-shot LLMs, plus a fine-tuned RoBERTa baseline, across two complementary benchmarks: GoEmotions and EmoBank, totaling 640,000 LLM responses. Zero-shot models diverge substantially from human distributions, and in-domain fine-tuning, not model scale, is required to close the gap. We formalize a lexical-grounding gradient through a quantitative transparency score that predicts per-category human--LLM agreement: LLMs reliably capture emotions with explicit lexical markers but systematically fail on pragmatically complex emotions requiring contextual inference, a pattern that replicates across both categorical and continuous emotion frameworks. We further propose three lightweight post-hoc calibration methods that reduce the distributional gap by up to 14\%, and provide actionable guidelines for when LLM emotion annotations can, and cannot, substitute for human labeling.

cs.CL

Multi-Stage Evolutionary Model Merging with Meta Data Driven Curriculum Learning for Sentiment-Specialized Large Language Modeling

The emergence of large language models (LLMs) has significantly transformed natural language processing (NLP), enabling more generalized models to perform various tasks with minimal training. However, traditional sentiment analysis methods, which focus on individual tasks such as sentiment classification or aspect-based analysis, are not practical for real-world applications that usually require handling multiple tasks. While offering flexibility, LLMs in sentiment-specific tasks often fall short of the required accuracy. Techniques like fine-tuning and evolutionary model merging help integrate models into a unified framework, which can improve the learning performance while reducing computational costs. The use of task meta-data and curriculum learning to optimize learning processes remains underexplored, while sentiment analysis is a critical task in NLP that requires high accuracy and scalability across multiple subtasks. In this study, we propose a hybrid learning model called Multi-stage Evolutionary Model Merging with Meta data driven Curriculum Learning (MEM-MCL), to enhance the sentiment analysis in large language modeling. In particular, expert models are created through instruction tuning for specific sentiment tasks and then merged using evolutionary algorithms to form a unified model. The merging process is optimized with weak data to enhance performance across tasks. The curriculum learning is incorporated to provide a learning sequence based on task difficulty, improving knowledge extraction from LLMs. Experiment results demonstrate that the proposed MEM-MCL model outperforms conventional LLMs in a majority of sentiment analysis tasks, achieving superior results across various subtasks.

cs.CL

CIEGAD: Cluster-Conditioned Interpolative and Extrapolative Framework for Geometry-Aware and Domain-Aligned Data Augmentation

In practical deep learning deployment, the scarcity of data and the imbalance of label distributions often lead to semantically uncovered regions within the real-world data distribution, hindering model training and causing misclassification near class boundaries as well as unstable behaviors in peripheral areas. Although recent large language models (LLMs) show promise for data augmentation, an integrated framework that simultaneously achieves directional control of generation, domain alignment, and quality control has not yet been fully established. To address these challenges, we propose a Cluster-conditioned Interpolative and Extrapolative framework for Geometry-Aware and Domain-aligned data augmentation (CIEGAD), which systematically complements both in-distribution and out-of-distribution semantically uncovered regions. CIEGAD constructs domain profiles through cluster conditioning, allocates generation with a hierarchical frequency-geometric allocation integrating class frequency and geometric indicators, and finely controls generation directions via the coexistence of interpolative and extrapolative synthesis. It further performs quality control through geometry-constrained filtering combined with an LLM-as-a-Judge mechanism. Experiments on multiple classification tasks demonstrate that CIEGAD effectively extends the periphery of real-world data distributions while maintaining high alignment between generated and real-world data as well as semantic diversity. In particular, for long-tailed and multi-class classification tasks, CIEGAD consistently improves F1 and recall, validating the triple harmony of distributional consistency, diversity, and quality. These results indicate that CIEGAD serves as a practically oriented data augmentation framework that complements underrepresented regions while preserving alignment with real-world data.

cs.LG

133-Tbps 1040-km (13$\times$80 km) Lumped-Amplified Transmission Over 22 THz in S-to-U-Band Using Hybrid Multiband Repeater with PPLN-Based Optical Parametric Amplifiers and EDFAs

We demonstrated 22.05-THz four-band long-haul transmission with a S-to-U-band lumped repeater consisting of PPLN-based optical parametric amplifiers and EDFAs over an 80-km-span SMF link. The achieved net bitrate was 133.06 Tbps at 1040 km with the 25.5-dBm fibre launch power designed by accounting for ISRS.

physics.optics

Programmable Photonic Unitary Processor Enables Parametrized Differentiable Long-Haul Spatial Division Multiplexed Transmission

The explosive growth of global data traffic demands scalable and energy-efficient optical communication systems. Spatial division multiplexing (SDM) using multicore or multimode fibers is a promising solution to overcome the capacity limit of single-mode fibers. However, long-haul SDM transmission faces significant challenges due to modal dispersion, which imposes heavy computational loads on digital signal processing (DSP) for signal equalization. Here, we propose parameterized SDM transmission, where programmable photonic unitary processors are installed at intermediate nodes. Instead of relying on conventional digital equalization only on the receiver side, our approach enables direct optimization of the SDM transmission channel itself by the programmable unitary processor, which reduces digital post-processing loads. We introduce a gradient-based optimization algorithm using a differentiable SDM transmission model to determine the optimal unitary transformation. As a key enabler, we first implemented telecom-grade programmable photonic unitary processor, achieving a low-loss (2.1 dB fiber-to-fiber), wideband (full C-band), polarization-independent, and high-fidelity (R2>96% across the C-band) operation. We experimentally demonstrate 1300-km transmission using a three-mode fiber, achieving strong agreement between simulation and experiment. The optimized photonic processor significantly reduces modal dispersion and post-processing complexity. Our results establish a scalable framework for integrating photonic computation into the optical layer, enabling more efficient, high-capacity optical networks.

physics.optics

Compressive dual-comb spectroscopy

Broadband, high resolution and rapid measurement of dual-comb spectroscopy (DCS) generates a large amount of data stream. We numerically demonstrate significant data compression of DCS spectra by using a compressive sensing technique. Our numerical simulation shows a compression rate of more than 100 with 3% error in mole fraction estimation of mid-infrared (MIR) DCS of two molecular species in a broadband (~30 THz) and high resolution (~115 MHz) condition. We also numerically demonstrate a massively parallel MIR DCS spectrum of 10 different molecular species can be reconstructed with a compression rate of 10.5 with a transmittance error of 0.003 from the original spectrum.

eess.SP

Time-stretch infrared spectroscopy

Improving spectral acquisition rate of broadband mid-infrared spectroscopy promises further advancements of molecular science and technology. Unlike the pump-probe spectroscopy that requires repeated measurements with different pump-probe delays, continuous spectroscopy running at a high spectral acquisition rate enables transient measurements of rapidly changing non-repeating phenomena or statistical analysis of a large amount of spectral data acquired within a short time. Recently, Fourier-transform infrared spectrometers (FT-IR) with rapid delay scan mechanisms including dual-comb spectrometers have significantly improved the measurement rate up to ~1 MSpectra/s that is fundamentally limited by the signal-to-noise ratio. Here, we overcome the limit and demonstrate the fastest continuous broadband vibrational spectrometer running at 80 MSpectra/s by implementing wavelength-swept time-stretch spectroscopy technique in the mid-infrared region. Our proof-of-concept experiment of the time-stretch infrared spectroscopy (TS-IR) demonstrates broadband absorption spectroscopy of phenylacetylene from 4.4 to 4.9 μm (2040-2270 cm-1) at a resolution of 15 nm (7.7 cm-1) with a superior signal-to-noise ratio of 85 without averaging and a shot-to-shot fluctuation of 1.3%.

physics.optics

Complementary Vibrational Spectroscopy

Vibrational spectroscopy, comprised of infrared absorption and Raman scattering spectroscopy, is widely used for label-free optical sensing and imaging in various scientific and industrial fields. The group theory states that the two molecular spectroscopy methods are sensitive to vibrations categorized in different point groups and provide complementary vibrational spectra. Therefore, complete vibrational information cannot be acquired by a single spectroscopic device, which has impeded the full potential of vibrational spectroscopy. Here, we demonstrate simultaneous infrared absorption and Raman scattering spectroscopy that allows us to measure the complete broadband vibrational spectra in the molecular fingerprint region with a single instrument based on an ultrashort pulsed laser. The system is based on dual-modal Fourier-transform spectroscopy enabled by efficient use of nonlinear optical effects. Our proof-of-concept experiment demonstrates rapid, broadband and high spectral resolution measurements of complementary spectra of organic liquids for precise and accurate molecular analysis.

physics.chem-ph