Searcharxiv⌕ Search

arXiv subjects

Alexander Martin

Publications and source records attributed to Alexander Martin.

32 records · Page 2Linked to original sources

MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval

Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on matching descriptive but vague queries with small collections of professionally edited, English-centric videos. To address this gap, we introduce $\textbf{MultiVENT 2.0}$, a large-scale, multilingual event-centric video retrieval benchmark featuring a collection of more than 218,000 news videos and 3,906 queries targeting specific world events. These queries specifically target information found in the visual content, audio, embedded text, and text metadata of the videos, requiring systems leverage all these sources to succeed at the task. Preliminary results show that state-of-the-art vision-language models struggle significantly with this task, and while alternative approaches show promise, they are still insufficient to adequately address this problem. These findings underscore the need for more robust multimodal retrieval systems, as effective video retrieval is a crucial step towards multimodal content understanding and generation.

cs.CV↗

Cross-Document Event-Keyed Summarization

Event-keyed summarization (EKS) requires summarizing a specific event described in a document given the document text and an event representation extracted from it. In this work, we extend EKS to the cross-document setting (CDEKS), in which summaries must synthesize information from accounts of the same event as given by multiple sources. We introduce SEAMUS (Summaries of Events Across Multiple Sources), a high-quality dataset for CDEKS based on an expert reannotation of the FAMUS dataset for cross-document argument extraction. We present a suite of baselines on SEAMUS -- covering both smaller, fine-tuned models, as well as zero- and few-shot prompted LLMs -- along with detailed ablations and a human evaluation study, showing SEAMUS to be a valuable benchmark for this new task.

cs.CL↗

Grounding Partially-Defined Events in Multimodal Data

How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially observable events, visual data does not facilitate analogous methods and, consequently, introduces unique challenges in event understanding. With the growing prevalence of vision-capable AI agents, these systems must be able to model events from collections of unstructured video data. To tackle robust event modeling in multimodal settings, we introduce a multimodal formulation for partially-defined events and cast the extraction of these events as a three-stage span retrieval task. We propose a corresponding benchmark for this task, MultiVENT-G, that consists of 14.5 hours of densely annotated current event videos and 1,168 text documents, containing 22.8K labeled event-centric entities. We propose a collection of LLM-driven approaches to the task of multimodal event analysis, and evaluate them on MultiVENT-G. Results illustrate the challenges that abstract event understanding poses and demonstrates promise in event-centric video-language systems.

cs.CL↗

Event-Keyed Summarization

We introduce event-keyed summarization (EKS), a novel task that marries traditional summarization and document-level event extraction, with the goal of generating a contextualized summary for a specific event, given a document and an extracted event structure. We introduce a dataset for this task, MUCSUM, consisting of summaries of all events in the classic MUC-4 dataset, along with a set of baselines that comprises both pretrained LM standards in the summarization literature, as well as larger frontier models. We show that ablations that reduce EKS to traditional summarization or structure-to-text yield inferior summaries of target events and that MUCSUM is a robust benchmark for this task. Lastly, we conduct a human evaluation of both reference and model summaries, and provide some detailed analysis of the results.

cs.CL↗

FAMuS: Frames Across Multiple Sources

Understanding event descriptions is a central aspect of language processing, but current approaches focus overwhelmingly on single sentences or documents. Aggregating information about an event \emph{across documents} can offer a much richer understanding. To this end, we present FAMuS, a new corpus of Wikipedia passages that \emph{report} on some event, paired with underlying, genre-diverse (non-Wikipedia) \emph{source} articles for the same event. Events and (cross-sentence) arguments in both report and source are annotated against FrameNet, providing broad coverage of different event types. We present results on two key event understanding tasks enabled by FAMuS: \emph{source validation} -- determining whether a document is a valid source for a target report event -- and \emph{cross-document argument extraction} -- full-document argument extraction for a target event from both its report and the correct source article. We release both FAMuS and our models to support further research.

cs.CL↗

Jurassic World Remake: Bringing Ancient Fossils Back to Life via Zero-Shot Long Image-to-Image Translation

With a strong understanding of the target domain from natural language, we produce promising results in translating across large domain gaps and bringing skeletons back to life. In this work, we use text-guided latent diffusion models for zero-shot image-to-image translation (I2I) across large domain gaps (longI2I), where large amounts of new visual features and new geometry need to be generated to enter the target domain. Being able to perform translations across large domain gaps has a wide variety of real-world applications in criminology, astrology, environmental conservation, and paleontology. In this work, we introduce a new task Skull2Animal for translating between skulls and living animals. On this task, we find that unguided Generative Adversarial Networks (GANs) are not capable of translating across large domain gaps. Instead of these traditional I2I methods, we explore the use of guided diffusion and image editing models and provide a new benchmark model, Revive-2I, capable of performing zero-shot I2I via text-prompting latent diffusion models. We find that guidance is necessary for longI2I because, to bridge the large domain gap, prior knowledge about the target domain is needed. In addition, we find that prompting provides the best and most scalable information about the target domain as classifier-guided diffusion models require retraining for specific use cases and lack stronger constraints on the target domain because of the wide variety of images they are trained on.

cs.CV↗

MegaWika: Millions of reports and their sources across 50 diverse languages

To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages, along with their 71 million referenced source materials. We process this dataset for a myriad of applications, going beyond the initial Wikipedia citation extraction and web scraping of content, including translating non-English articles for cross-lingual applications and providing FrameNet parses for automated semantic analysis. MegaWika is the largest resource for sentence-level report generation and the only report generation dataset that is multilingual. We manually analyze the quality of this resource through a semantically stratified sample. Finally, we provide baseline results and trained models for crucial steps in automated report generation: cross-lingual question answering and citation retrieval.

cs.CL↗

Coherent dynamics in a five-level atomic system

The coherent control of multi-partite quantum systems presents one of the central prerequisites in state-of-the-art quantum information processing. With the added benefit of inherent high-fidelity detection capability, atomic quantum systems in high-energy internal states, such as metastable noble gas atoms, promote themselves as ideal candidates for advancing quantum science in fundamental aspects and technological applications. Using laser-cooled neon atoms in the metastable $^3$P$_2$ state of state $1s^2 2s^2 2p^5 3s$ (LS-coupling notation) (Racah notation: $^2P_{3/2}\,3s[3/2]_2$) with five $m_F$-sublevels, experimental methods for the preparation of all Zeeman sublevels |m_J> = |+2>, |+1>, |0>, |-1>, |-2> as well as the coherent control of superposition states in the five-level system |+2> ... |-2>, in the three-level system |+2>, |+1>, |0>, and in the two-level system |+2>, |+1> are presented. The methods are based on optimized radio frequency and laser pulse sequences. The state evolution is described with a simple, semiclassical model. The coherence properties of the prepared states are studied using Ramsey and spin echo measurements.

quant-ph↗

The Bipartite Boolean Quadric Polytope with Multiple-Choice Constraints

We consider the bipartite boolean quadric polytope (BQP) with multiple-choice constraints and analyse its combinatorial properties. The well-studied BQP is defined as the convex hull of all quadric incidence vectors over a bipartite graph. In this work, we study the case where there is a partition on one of the two bipartite node sets such that at most one node per subset of the partition can be chosen. This polytope arises, for instance, in pooling problems with fixed proportions of the inputs at each pool. We show that it inherits many characteristics from BQP, among them a wide range of facet classes and operations which are facet preserving. Moreover, we characterize various cases in which the polytope is completely described via the relaxation-linearization inequalities. The special structure induced by the additional multiple-choice constraints also allows for new facet-preserving symmetries as well as lifting operations. Furthermore, it leads to several novel facet classes as well as extensions of these via lifting. We additionally give computationally tractable exact separation algorithms, most of which run in polynomial time. Finally, we demonstrate the strength of both the inherited and the new facet classes in computational experiments on random as well as real-world problem instances. It turns out that in many cases we can close the optimality gap almost completely via cutting planes alone, and, consequently, solution times can be reduced significantly.

math.OC↗

An Online-Learning Approach to Inverse Optimization

In this paper, we demonstrate how to learn the objective function of a decision-maker while only observing the problem input data and the decision-maker's corresponding decisions over multiple rounds. We present exact algorithms for this online version of inverse optimization which converge at a rate of $ \mathcal{O}(1/\sqrt{T}) $ in the number of observations~$T$ and compare their further properties. Especially, they all allow taking decisions which are essentially as good as those of the observed decision-maker already after relatively few iterations, but are suited best for different settings each. Our approach is based on online learning and works for linear objectives over arbitrary feasible sets for which we have a linear optimization oracle. As such, it generalizes previous approaches based on KKT-system decomposition and dualization. We also introduce several generalizations, such as the approximate learning of non-linear objective functions, dynamically changing as well as parameterized objectives and the case of suboptimal observed decisions. When applied to the stochastic offline case, our algorithms are able to give guarantees on the quality of the learned objectives in expectation. Finally, we show the effectiveness and possible applications of our methods in indicative computational experiments.

math.OC↗

Pricing and clearing combinatorial markets with singleton and swap orders: Efficient algorithms for the futures opening auction problem

In this article we consider combinatorial markets with valuations only for singletons and pairs of buy/sell-orders for swapping two items in equal quantity. We provide an algorithm that permits polynomial time market-clearing and -pricing. The results are presented in the context of our main application: the futures opening auction problem. Futures contracts are an important tool to mitigate market risk and counterparty credit risk. In futures markets these contracts can be traded with varying expiration dates and underlyings. A common hedging strategy is to roll positions forward into the next expiration date, however this strategy comes with significant operational risk. To address this risk, exchanges started to offer so-called futures contract combinations, which allow the traders for swapping two futures contracts with different expiration dates or for swapping two futures contracts with different underlyings. In theory, the price is in both cases the difference of the two involved futures contracts. However, in particular in the opening auctions price inefficiencies often occur due to suboptimal clearing, leading to potential arbitrage opportunities. We present a minimum cost flow formulation of the futures opening auction problem that guarantees consistent prices. The core ideas are to model orders as arcs in a network, to enforce the equilibrium conditions with the help of two hierarchical objectives, and to combine these objectives into a single weighted objective while preserving the price information of dual optimal solutions. The resulting optimization problem can be solved in polynomial time and computational tests establish an empirical performance suitable for production environments.

math.OC↗

General framework for treating generation, propagation, and polarization of luminescence in anisotropic media

Complete polarimeters deliver the full polarization transfer matrix of a medium that relates input polarization states to output polarization states. In order to interpret the Mueller matrix of a luminescent medium at the emission frequency, accountings are required for polarization transformations of the medium at the excitation frequency, the light scattering event, and the polarization transformations at the emission frequency. A general framework for this kind of analysis is presented herein. The fluorescence Mueller matrix is expressed as a product of Mueller matrices of the medium at the excitation and emission frequencies and a scattering matrix, integrated over path length. The Stokes vector for the incident light, evolving according to the Mueller matrix of the medium at the excitation frequency, is multiplied by the scattering matrix to give a Stokes vector of the emitted light along the propagation direction. The scattering matrix is derived from an incoherent orientation ensemble average over a phenomenological scattering tensor that embodies intrinsic molecular scattering proprieties, and dynamical processes that occur during the excited state. The general framework is evaluated for three anisotropic materials carrying luminescent dye molecules including the following: a chiral fluid, a stretched polymer film, and a chiral, biaxial crystal. In the latter case, the most complex, the Mueller matrix was collected in conoscopic illumination and the fluorescence Mueller matrix, mapped in k-space, was fully simulated by the strategy outlined above. Luminescence spectroscopy has typically stood apart from the developments in polarimetry of the past two generations. This need not be indefinite.

physics.optics↗

External Cavity Diode Laser Setup with Two Interference Filters

We present an external cavity diode laser setup using two identical, commercially available interference filters operated in the blue wavelength range around 450 nm. The combination of the two filters decreases the transmission width, while increasing the edge steepness without a significant reduction in peak transmittance. Due to the broad spectral transmission of such interference filters compared to the internal mode spacing of blue laser diodes, an additional locking scheme, based on Hänsch-Couillaud locking to a cavity, has been added to improve the stability. The laser is stabilized to a line in the tellurium spectrum via saturation spectroscopy, and single-frequency operation for a duration of two days is demonstrated by monitoring the error signal of the lock and the piezo drive compensating the length change of the external resonator due to air pressure variations. Additionally, transmission curves of the filters and the spectra of a sample of diodes are given.

physics.optics↗

Strict linear prices in non-convex European day-ahead electricity markets

The European power grid can be divided into several market areas where the price of electricity is determined in a day-ahead auction. Market participants can provide continuous hourly bid curves and combinatorial bids with associated quantities given the prices. The goal of our auction is to maximize the economic surplus of all participants subject to quantity constraints and price constraints. The price constraints ensure that no one incurs a loss. Only traders who submitted a combinatorial bid might miss a not-realized profit. The resulting problem is a large scale mathematical program with equilibrium constraints (MPEC) and binary variables that cannot be solved efficiently by standard solvers. We present an exact algorithm and a fast heuristic for this type of problem. Both algorithms decompose the MPEC into a master problem (a MIQP) and pricing subproblems (LPs). The modeling technique and the algorithms are applicable to a wide variety of combinatorial auctions that are based on mixed integer programs.

math.OC↗