SearcharxivSearch

arXiv subjects

Luca Cultrera

Publications and source records attributed to Luca Cultrera.

At least 19 recordsLinked to original sources

CW SRF Gun generating beam parameters sufficient for CW hard-X-ray FEL

SRF CW accelerator constructed for Coherent electron Cooling (CeC) Proof-of-principle (POP) experiment at Brookhaven National Laboratory has frequently demonstrated record parameters using 1.5 nC 350 ps long electron bunches, typically compressed to FWHM of 30 ps using ballistic compression. We report experimental demonstration of CW electron beam with parameters fully satisfying requirements for hard X-ray FEL and significantly exceeding those demonstrated by APEX LCLS II electron gun. This was achieved using a 10-year-old SRF gun with a modest accelerating gradient of $\sim$15 MV/m, a bunching cavity followed by ballistic compression to generate 100 pC, $\sim$15 ps FWHM electron bunches with a normalized slice emittance of $\sim$0.2 mm-mrad and a normalized projected emittance of $\sim$0.25 mm-mrad. Hence, in this paper, we present an alternative method for generating CW electron beams for hard-X-ray FELs using existing and proven accelerator technology. We present a description of the accelerator system settings, details of projected and slice emittance measurements as well as relevant beam dynamics simulations.

physics.acc-ph

Drone Detection with Event Cameras

The diffusion of drones presents significant security and safety challenges. Traditional surveillance systems, particularly conventional frame-based cameras, struggle to reliably detect these targets due to their small size, high agility, and the resulting motion blur and poor performance in challenging lighting conditions. This paper surveys the emerging field of event-based vision as a robust solution to these problems. Event cameras virtually eliminate motion blur and enable consistent detection in extreme lighting. Their sparse, asynchronous output suppresses static backgrounds, enabling low-latency focus on motion cues. We review the state-of-the-art in event-based drone detection, from data representation methods to advanced processing pipelines using spiking neural networks. The discussion extends beyond simple detection to cover more sophisticated tasks such as real-time tracking, trajectory forecasting, and unique identification through propeller signature analysis. By examining current methodologies, available datasets, and the distinct advantages of the technology, this work demonstrates that event-based vision provides a powerful foundation for the next generation of reliable, low-latency, and efficient counter-UAV systems.

cs.CV

Spike-TBR: a Noise Resilient Neuromorphic Event Representation

Event cameras offer significant advantages over traditional frame-based sensors, including higher temporal resolution, lower latency and dynamic range. However, efficiently converting event streams into formats compatible with standard computer vision pipelines remains a challenging problem, particularly in the presence of noise. In this paper, we propose Spike-TBR, a novel event-based encoding strategy based on Temporal Binary Representation (TBR), addressing its vulnerability to noise by integrating spiking neurons. Spike-TBR combines the frame-based advantages of TBR with the noise-filtering capabilities of spiking neural networks, creating a more robust representation of event streams. We evaluate four variants of Spike-TBR, each using different spiking neurons, across multiple datasets, demonstrating superior performance in noise-affected scenarios while improving the results on clean data. Our method bridges the gap between spike-based and frame-based processing, offering a simple noise-resilient solution for event-driven vision applications.

cs.CV

Spatio-temporal Transformers for Action Unit Classification with Event Cameras

Face analysis has been studied from different angles to infer emotion, poses, shapes, and landmarks. Traditionally RGB cameras are used, yet for fine-grained tasks standard sensors might not be up to the task due to their latency, making it impossible to record and detect micro-movements that carry a highly informative signal, which is necessary for inferring the true emotions of a subject. Event cameras have been increasingly gaining interest as a possible solution to this and similar high-frame rate tasks. We propose a novel spatiotemporal Vision Transformer model that uses Shifted Patch Tokenization (SPT) and Locality Self-Attention (LSA) to enhance the accuracy of Action Unit classification from event streams. We also address the lack of labeled event data in the literature, which can be considered one of the main causes of an existing gap between the maturity of RGB and neuromorphic vision models. Gathering data is harder in the event domain since it cannot be crawled from the web and labeling frames should take into account event aggregation rates and the fact that static parts might not be visible in certain frames. To this end, we present FACEMORPHIC, a temporally synchronized multimodal face dataset composed of RGB videos and event streams. The dataset is annotated at a video level with facial Action Units and contains streams collected with various possible applications, ranging from 3D shape estimation to lip-reading. We then show how temporal synchronization can allow effective neuromorphic face analysis without the need to manually annotate videos: we instead leverage cross-modal supervision bridging the domain gap by representing face shapes in a 3D space. Our proposed model outperforms baseline methods by effectively capturing spatial and temporal information, crucial for recognizing subtle facial micro-expressions.

cs.CV

Garment Attribute Manipulation with Multi-level Attention

In the rapidly evolving field of online fashion shopping, the need for more personalized and interactive image retrieval systems has become paramount. Existing methods often struggle with precisely manipulating specific garment attributes without inadvertently affecting others. To address this challenge, we propose GAMMA (Garment Attribute Manipulation with Multi-level Attention), a novel framework that integrates attribute-disentangled representations with a multi-stage attention-based architecture. GAMMA enables targeted manipulation of fashion image attributes, allowing users to refine their searches with high accuracy. By leveraging a dual-encoder Transformer and memory block, our model achieves state-of-the-art performance on popular datasets like Shopping100k and DeepFashion.

cs.CV

Neuromorphic Facial Analysis with Cross-Modal Supervision

Traditional approaches for analyzing RGB frames are capable of providing a fine-grained understanding of a face from different angles by inferring emotions, poses, shapes, landmarks. However, when it comes to subtle movements standard RGB cameras might fall behind due to their latency, making it hard to detect micro-movements that carry highly informative cues to infer the true emotions of a subject. To address this issue, the usage of event cameras to analyze faces is gaining increasing interest. Nonetheless, all the expertise matured for RGB processing is not directly transferrable to neuromorphic data due to a strong domain shift and intrinsic differences in how data is represented. The lack of labeled data can be considered one of the main causes of this gap, yet gathering data is harder in the event domain since it cannot be crawled from the web and labeling frames should take into account event aggregation rates and the fact that static parts might not be visible in certain frames. In this paper, we first present FACEMORPHIC, a multimodal temporally synchronized face dataset comprising both RGB videos and event streams. The data is labeled at a video level with facial Action Units and also contains streams collected with a variety of applications in mind, ranging from 3D shape estimation to lip-reading. We then show how temporal synchronization can allow effective neuromorphic face analysis without the need to manually annotate videos: we instead leverage cross-modal supervision bridging the domain gap by representing face shapes in a 3D space.

cs.CV

Prompt and Prejudice

This paper investigates the impact of using first names in Large Language Models (LLMs) and Vision Language Models (VLMs), particularly when prompted with ethical decision-making tasks. We propose an approach that appends first names to ethically annotated text scenarios to reveal demographic biases in model outputs. Our study involves a curated list of more than 300 names representing diverse genders and ethnic backgrounds, tested across thousands of moral scenarios. Following the auditing methodologies from social sciences we propose a detailed analysis involving popular LLMs/VLMs to contribute to the field of responsible AI by emphasizing the importance of recognizing and mitigating biases in these systems. Furthermore, we introduce a novel benchmark, the Pratical Scenarios Benchmark (PSB), designed to assess the presence of biases involving gender or demographic prejudices in everyday decision-making scenarios as well as practical scenarios where an LLM might be used to make sensible decisions (e.g., granting mortgages or insurances). This benchmark allows for a comprehensive comparison of model behaviors across different demographic categories, highlighting the risks and biases that may arise in practical applications of LLMs and VLMs.

cs.CL

UV hybrid photon detector based on GaN photocathodes and Si low gain avalanche diode

Photon detectors featuring single-photon sensitivity play a crucial role in various scientific domains, including high-energy physics, astronomy, and quantum optics. Fast response time, high quantum efficiency, and minimal dark counts are the characteristics that render them ideal candidates for detecting individual photons with exceptional signal-to-noise ratios, at frequencies in the the range of hundreds of MHz. Here, we report on our first design and operational results on a Hybrid Photon Detector (HPD) that combines the high quantum efficiency of a Gallium Nitride (GaN) photocathode and the low noise characteristics of a Si-based Low-Gain Avalanche Diode (LGAD). This hybrid detection scheme has the potential to reach single-photon detection sensitivity with high quantum efficiency, low noise levels and capable of operating at hundreds of MHz repetition rates.

physics.ins-det

Neuromorphic Face Analysis: a Survey

Neuromorphic sensors, also known as event cameras, are a class of imaging devices mimicking the function of biological visual systems. Unlike traditional frame-based cameras, which capture fixed images at discrete intervals, neuromorphic sensors continuously generate events that represent changes in light intensity or motion in the visual field with high temporal resolution and low latency. These properties have proven to be interesting in modeling human faces, both from an effectiveness and a privacy-preserving point of view. Neuromorphic face analysis however is still a raw and unstructured field of research, with several attempts at addressing different tasks with no clear standard or benchmark. This survey paper presents a comprehensive overview of capabilities, challenges and emerging applications in the domain of neuromorphic face analysis, to outline promising directions and open issues. After discussing the fundamental working principles of neuromorphic vision and presenting an in-depth overview of the related research, we explore the current state of available data, standard data representations, emerging challenges, and limitations that require further investigation. This paper aims to highlight the recent process in this evolving field to provide to both experienced and newly come researchers an all-encompassing analysis of the state of the art along with its problems and shortcomings.

cs.CV

Neuromorphic Valence and Arousal Estimation

Recognizing faces and their underlying emotions is an important aspect of biometrics. In fact, estimating emotional states from faces has been tackled from several angles in the literature. In this paper, we follow the novel route of using neuromorphic data to predict valence and arousal values from faces. Due to the difficulty of gathering event-based annotated videos, we leverage an event camera simulator to create the neuromorphic counterpart of an existing RGB dataset. We demonstrate that not only training models on simulated data can still yield state-of-the-art results in valence-arousal estimation, but also that our trained models can be directly applied to real data without further training to address the downstream task of emotion recognition. In the paper we propose several alternative models to solve the task, both frame-based and video-based.

cs.CV

Addressing Limitations of State-Aware Imitation Learning for Autonomous Driving

Conditional Imitation learning is a common and effective approach to train autonomous driving agents. However, two issues limit the full potential of this approach: (i) the inertia problem, a special case of causal confusion where the agent mistakenly correlates low speed with no acceleration, and (ii) low correlation between offline and online performance due to the accumulation of small errors that brings the agent in a previously unseen state. Both issues are critical for state-aware models, yet informing the driving agent of its internal state as well as the state of the environment is of crucial importance. In this paper we propose a multi-task learning agent based on a multi-stage vision transformer with state token propagation. We feed the state of the vehicle along with the representation of the environment as a special token of the transformer and propagate it throughout the network. This allows us to tackle the aforementioned issues from different angles: guiding the driving policy with learned stop/go information, performing data augmentation directly on the state of the vehicle and visually explaining the model's decisions. We report a drastic decrease in inertia and a high correlation between offline and online metrics.

cs.CV

Neuromorphic Event-based Facial Expression Recognition

Recently, event cameras have shown large applicability in several computer vision fields especially concerning tasks that require high temporal resolution. In this work, we investigate the usage of such kind of data for emotion recognition by presenting NEFER, a dataset for Neuromorphic Event-based Facial Expression Recognition. NEFER is composed of paired RGB and event videos representing human faces labeled with the respective emotions and also annotated with face bounding boxes and facial landmarks. We detail the data acquisition process as well as providing a baseline method for RGB and event data. The collected data captures subtle micro-expressions, which are hard to spot with RGB data, yet emerge in the event domain. We report a double recognition accuracy for the event-based approach, proving the effectiveness of a neuromorphic approach for analyzing fast and hardly detectable expressions and the emotions they conceal.

cs.CV

Spin polarized electron beams production beyond III-V semiconductors

This paper summarizes the state of the art of photocathode based on III-V semiconductors for spin polarized electron beam production. The limitations preventing this class of material to provide the long term reliability at the highest average beam currents necessary for some of the new accelerator facilities or proposed upgrades of existing ones are illustrated. Promising alternative classes of materials are identified showing properties that can be leveraged to synthesize photocathode structures that can outperform III-V semiconductors in the production of spin polarized electron beams and support the operating conditions of advanced electron sources for new facilities.

physics.acc-ph

Operation of Cs-Sb-O activated GaAs in a high voltage DC electron gun at high average current

Negative Electron Affinity (NEA) activated GaAs photocathodes are the most popular option for generating a high current (> 1 mA) spin-polarized electron beam. Despite its popularity, a short operational lifetime is the main drawback of this material. Recent works have shown that the lifetime can be improved by using a robust Cs-Sb-O NEA layer with minimal adverse effects. In this work, we operate GaAs photocathodes with this new activation method in a high voltage environment to extract a high current. We observed spectral dependence on the lifetime improvement. In particular, we saw a 45% increase in the lifetime at 780 nm for Cs-Sb-O activated GaAs compared to Cs-O activated GaAs.

physics.acc-ph

Monte Carlo Modeling of Spin-polarized Photoemission from p-doped GaAs Activated to Negative Electron Affinity

The anticorrelation between quantum efficiency (QE) and electron spin polarization (ESP) from a p-doped GaAs activated to negative electron affinity (NEA) is studied in detail using an ensemble Monte Carlo approach. The photoabsorption, momentum and spin relaxation during transport, and tunnelling of electrons through the surface potential barrier are modeled to identify fundamental mechanisms, which limit the efficiency of GaAs spin-polarized electron sources. In particular, we study the response of QE and ESP to various parameters such as the photoexcitation energy, doping density, and electron affinity level. Our modeling results for various transport and emission characteristics are in a good agreement with available experimental data. Our findings show that the behaviour of both QE and ESP at room temperature can be fully explained by the bulk relaxation mechanisms and the time which electrons spend in the material before being emitted.

physics.app-ph

Modelling the Statistics of Cyclic Activities by Trajectory Analysis on the Manifold of Positive-Semi-Definite Matrices

In this paper, a model is presented to extract statistical summaries to characterize the repetition of a cyclic body action, for instance a gym exercise, for the purpose of checking the compliance of the observed action to a template one and highlighting the parts of the action that are not correctly executed (if any). The proposed system relies on a Riemannian metric to compute the distance between two poses in such a way that the geometry of the manifold where the pose descriptors lie is preserved; a model to detect the begin and end of each cycle; a model to temporally align the poses of different cycles so as to accurately estimate the \emph{cross-sectional} mean and variance of poses across different cycles. The proposed model is demonstrated using gym videos taken from the Internet.

cs.CV

Explaining Autonomous Driving by Learning End-to-End Visual Attention

Current deep learning based autonomous driving approaches yield impressive results also leading to in-production deployment in certain controlled scenarios. One of the most popular and fascinating approaches relies on learning vehicle controls directly from data perceived by sensors. This end-to-end learning paradigm can be applied both in classical supervised settings and using reinforcement learning. Nonetheless the main drawback of this approach as also in other learning problems is the lack of explainability. Indeed, a deep network will act as a black-box outputting predictions depending on previously seen driving patterns without giving any feedback on why such decisions were taken. While to obtain optimal performance it is not critical to obtain explainable outputs from a learned agent, especially in such a safety critical field, it is of paramount importance to understand how the network behaves. This is particularly relevant to interpret failures of such systems. In this work we propose to train an imitation learning based agent equipped with an attention model. The attention model allows us to understand what part of the image has been deemed most important. Interestingly, the use of attention also leads to superior performance in a standard benchmark using the CARLA driving simulator.

cs.CV

Improved lifetime of a high spin polarization superlattice photocathode

Negative Electron Affinity (NEA) activated surfaces are required to extract highly spin polarized electron beams from GaAs-based photocathodes, but they suffer extreme sensitivity to poor vacuum conditions that results in rapid degradation of quantum efficiency. We report on series of unconventional NEA activations on surfaces of bulk GaAs with Cs, Sb, and O2 using different methods of oxygen exposure for optimizing photocathode performance. One order of magnitude improvement in lifetime with respect to the standard Cs-O2 activation is achieved without significant loss on electron spin polarization and quantum efficiency by codepositing Cs, Sb, and O2. A strained GaAs/GaAsP superlattice sample activated with the codeposition method demonstrated similar enhancement in lifetime near the photoemission threshold while maintaining 90% spin polarization.

physics.app-ph