SearcharxivSearch

arXiv subjects

Yeeun Choi

Publications and source records attributed to Yeeun Choi.

7 recordsLinked to original sources

Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding

When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction to pre-model high-level relations in videos, despite not knowing the downstream query at build time. We instead prioritize high-recall retrievability during memory building, and defer query-specific, high-level relation composition to inference time. To this end, we propose MERIT(Multi-key Episodic Retrieval with Inference-time Temporal expansion), a simple yet effective agentic framework for ultra-long video understanding. First, we formulate an episodic multi-key representation that enables precise retrieval of fine-grained memories through a simple key-matching mechanism. Second, we introduce a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction. This is achieved by expanding the temporal scope exclusively around the retrieved segments at inference time. By leveraging simple key-matching with this on-demand temporal expansion, MERIT achieves state-of-the-art performance across three long-video benchmarks: EgoLifeQA, LVBench, and Video-MME (Long).

cs.CV

SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and context. Existing testbeds rely on a few expert-authored domains, vary mainly strategic posture, and score every turn against every topic, introducing off-topic noise. We introduce SoCRATES, a benchmark for evaluating proactive LLM mediators in realistic, multi-domain testbeds. It constructs scenarios from real conflicts through an agentic pipeline across eight domains, probes five socio-cognitive adaptation axes (strategic posture, party composition, history length, emotional reactivity, and cultural identity), and scores each topic only on the turns that advance it via a topic-localized evaluator. The evaluator reaches 0.82 alignment with human experts, more than doubling a per-turn baseline. Benchmarking eight frontier LLMs, we find that even the strongest mediator closes only about a third of the unmediated consensus gap under diverse and realistic testbeds, with performance varying sharply by socio-cognitive axis, highlighting that progress lies in social adaptation to diverse conditions.

cs.AI

Nonlinear Electro-Optic Visible Photonic Circuits for Solid-State Quantum Defects

Integrated visible photonic engines for solid-state quantum defects provide a foundation for scalable quantum networks. While miniaturization is advancing, active manipulation remains limited by the difficulty of achieving simultaneous milliwatt-scale visible light generation and high-contrast modulation. Despite extensive efforts, the concurrent chip-scale realization of nonlinear frequency conversion and fast temporal gating for high-fidelity quantum control has remained elusive. Here, we demonstrate a monolithic thin-film lithium niobate (TFLN) platform integrating periodically poled frequency conversion with GHz-bandwidth electro-optic (EO) switching. The device delivers off-chip green-light power exceeding 1 mW with an extinction ratio (ER) of 42.2 dB, enabling coherent spin control and time-resolved lifetime measurements of individual nitrogen-vacancy (NV) centers in diamond through nanosecond gating. System performance is validated through pulsed optically detected magnetic resonance (ODMR), Rabi oscillations, and Ramsey interference, supported by time-tagged photon counting with nanosecond resolution. By unifying sufficient nonlinear light generation with high-speed active manipulation, this platform establishes a scalable framework for the realization of high-rate quantum communication nodes.

physics.optics

Self-Aligned Heterogeneous Quantum Photonic Integration

Integrated quantum photonics holds significant promise for scalable photonic quantum information processing, quantum repeaters, and quantum networks, but its development is hindered by the mismatch between materials hosting high-quality quantum emitters and those compatible with mature photonic technologies. Heterogeneous integration offers a potential solution to this challenge, yet practical implementations have been limited by inevitable insertion losses at material interfaces. Here, we present a self-aligned heterogeneous quantum photonic integration approach that can deterministically achieve near-unity coupling efficiency at the interface. To showcase our approach, we demonstrate Purcell enhancement of a silicon vacancy (SiV) center in diamond induced by a heterogeneous photonic crystal cavity defined by titanium dioxide (TiO2), as well as optical spin control and readout via a TiO2 photonic circuit. We further show that, when combined with inverse photonic design, our approach enables efficient and broadband collection of single photons from a color center into a heterogeneous waveguide. Our approach is not restricted to SiV centers or TiO2; it can be broadly applied to integrate diverse solid-state quantum emitters with thin-film photonic devices where conformal deposition is possible. Together, these results establish a practical route to scalable quantum photonic integrated circuits that combine high-quality quantum emitters with technologically mature photonic platforms.

physics.optics

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

We introduce Hierarchical Streaming Video Understanding, a task that combines online temporal action localization with free-form description generation. Given the scarcity of datasets with hierarchical and fine-grained temporal annotations, we demonstrate that LLMs can effectively group atomic actions into higher-level events, enriching existing datasets. We then propose OpenHOUSE (Open-ended Hierarchical Online Understanding System for Events), which extends streaming action perception beyond action classification. OpenHOUSE features a specialized streaming module that accurately detects boundaries between closely adjacent actions, nearly doubling the performance of direct extensions of existing methods. We envision the future of streaming action perception in the integration of powerful generative models, with OpenHOUSE representing a key step in that direction.

cs.CV

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in vision-language understanding, reasoning, and generation. However, they struggle with tasks requiring fine-grained localization and reasoning in high-resolution images. This constraint stems from the fact that MLLMs are fine-tuned with fixed image resolution to align with the pre-trained image encoder used in MLLM. Consequently, feeding high-resolution images directly into MLLMs leads to poor generalization due to a train-test resolution discrepancy, while downsampling these images-although ensuring consistency-compromises fine-grained visual details and ultimately degrades performance. To address this challenge, we propose Extract Candidate then Predict (ECP), a novel training-free, task-agnostic two-stage framework designed to enhance MLLM performance on high-resolution images. The key intuition behind ECP is that while MLLMs struggle with high-resolution images, their predictions on downsampled images still contain implicit localization cues. By first identifying candidate region using the coarse prediction and then predicting the final output based on candidate region, ECP effectively preserves fine-grained details while mitigating the challenges posed by high-resolution data. We validate our framework on 4K GUI grounding and 4K, 8K MLLM perception, achieving +21.3%, +5.8%, +5.2% absolute improvement compared to baseline respectively, demonstrating its effectiveness. Code is available at https://github.com/yenncye/ECP.

cs.CV

Diamond molecular balance: Revolutionizing high-resolution mass spectrometry from MDa to TDa at room temperature

The significance of mass spectrometry lies in its unparalleled ability to accurately identify and quantify molecules in complex samples, providing invaluable insights into molecular structures and interactions. Here, we leverage diamond nanostructures as highly sensitive mass sensors by utilizing a self-excitation mechanism under an electron beam in a conventional scanning electron microscope (SEM). The diamond molecular balance (DMB) exhibits an exceptional mass resolution of 0.36 MDa, based on its outstanding mechanical quality factor and frequency stability, along with an extensive dynamic range from MDa to TDa. This positions the DMB at the forefront of molecular balances operating at room temperature. Notably, the DMB demonstrates its ability to measure the mass of a single bacteriophage T4 by precisely locating the analyte on the device. These findings highlight the groundbreaking potential of the DMB as a revolutionary tool for mass spectrometry at room temperature.

cond-mat.mes-hall