SearcharxivSearch

arXiv subjects

Sahar Froim

Publications and source records attributed to Sahar Froim.

5 recordsLinked to original sources

When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization

Annotators routinely disagree on laughter boundaries and subtle chuckles, yet temporal laughter localization typically evaluates against a single reference annotation. We show that this disagreement is structured rather than random noise. Re-annotating the SMILE-Temporal benchmark (672 videos, 1,683 events) with 3-5 annotators per video (alpha = 0.757), we find systematic patterns: disagreement is 1.73 times larger at offsets than onsets, far more common for chuckles than full laughs (77% vs. 20%), and predictable from event attributes (AUC = 0.831). Evaluating against a single annotator breaks down under this structure: system scores shift by 0.246 F1 depending on the chosen ground truth, correctly ranking systems only 69.7% of the time (vs. 80% against all annotators). We propose a disagreement-calibrated evaluation that scores predictions against the full annotator distribution using conformally calibrated tolerance bands (wider at offsets, 0.727s, than onsets, 0.5s). The per-annotator annotations and analysis code are available at https://github.com/WSCSports/MTLLFM-temporal-laughter-localization.

cs.CV

MTLLFM: Multimodal-Temporal Laughter Localization: UR-FUNNY-Temporal and SMILE-Temporal Benchmarks with an Adaptive Multimodal Fusion Model

Detecting laughter in video is essential for affective computing and narrative understanding, yet existing approaches treat it as coarse clip-level classification, failing to capture precise temporal boundaries of brief, transient laughter events. We address this gap with two complementary contributions. First, we introduce UR-FUNNY-Temporal and SMILE-Temporal, fully annotated temporal laughter datasets extending two widely-used humor benchmarks. Our annotations cover over 11,053 videos (78.8 hours) and provide precise onset/offset boundaries for each laughter event, along with rich metadata distinguishing speaker vs. audience laughter, modality dominance (acoustic, visual, or both), and intensity levels. Second, we propose a lightweight weakly-supervised framework for temporal laughter localization. Our architecture combines fixed HuBERT and MAE encoders with temporal softmax pooling and adaptive modality gating, learning fine-grained temporal grounding from clip-level labels without requiring frame-level annotations during training. Experiments across three datasets demonstrate that our approach substantially outperforms multimodal foundation models including Gemini 3 Flash, achieving 99% F1 and 68.1% localization precision on sports broadcast data. Ablations validate each architectural component. Furthermore, our precise temporal tags improve downstream laughter reasoning by 227% on CIDEr, enabling GPT-3.5 to outperform GPT-4o. The code, UR-FUNNY-Temporal and SMILE-Temporal datasets are publicly available at https://github.com/WSCSports/MTLLFM-temporal-laughter-localization.

cs.CV

Beam Profiler Network (BPNet) -- A Deep Learning Approach to Mode Demultiplexing of Laguerre-Gaussian Optical Beams

The transverse field profile of light is being recognized as a resource for classical and quantum communications for which reliable methods of sorting or demultiplexing spatial optical modes are required. Here, we demonstrate, experimentally, state-of-the-art mode demultiplexing of Laguerre-Gaussian beams according to both their orbital angular momentum and radial topological numbers using a flow of two concatenated deep neural networks. The first network serves as a transfer function from experimentally-generated to ideal numerically-generated data, while using a unique "Histogram Weighted Loss" function that solves the problem of images with limited significant information. The second network acts as a spatial-modes classifier. Our method uses only the intensity profile of modes or their superposition, making the phase information redundant.

eess.IV

Particle trapping and conveying using an optical Archimedes' screw

Trapping and manipulation of particles using laser beams has become an important tool in diverse fields of research. In recent years, particular interest is given to the problem of conveying optically trapped particles over extended distances either down or upstream the direction of the photons momentum flow. Here, we propose and demonstrate experimentally an optical analogue of the famous Archimedes' screw where the rotation of a helical-intensity beam is transferred to the axial motion of optically-trapped micro-meter scale airborne carbon based particles. With this optical screw, particles were easily conveyed with controlled velocity and direction, upstream or downstream the optical flow, over a distance of half a centimeter. Our results offer a very simple optical conveyor that could be adapted to a wide range of optical trapping scenarios.

physics.optics

Breaking the temporal resolution limit by superoscillating optical beats

Band-limited functions can oscillate locally at an arbitrarily fast rate through an interference phenomenon known as superoscillations. Using an optical pulse with a superoscillatory envelope we experimentally break the temporal Fourier-transform limit having a temporal feature which is approximately three times shorter than the duration of a transform-limited Gaussian pulse having a comparable bandwidth while maintaining $29.5\%$ visibility. Numerical simulations demonstrate the ability of such signals to achieve temporal super-resolution.

physics.optics