SearcharxivSearch

arXiv subjects

Ahmad Rahimi

Publications and source records attributed to Ahmad Rahimi.

12 recordsLinked to original sources

Deterministic positioning of circular Bragg gratings using atomic force lithography for high-performance quantum dot light sources

Semiconductor quantum dots (QDs) grown by molecular beam epitaxy are excellent quantum emitters, but their random spatial distribution hinders deterministic coupling to optical microcavities. We demonstrate a room-temperature atomic force microscopy (AFM)-assisted nano-oxidation lithography technique enabling QD positioning with a radial displacement of $51(28)$ nm. Free-standing asymmetric circular Bragg gratings incorporating AFM-positioned GaAs QDs exhibit a $245$-fold photoluminescence enhancement and fine-structure splitting (FSS) comparable to bulk QDs. Polarization-resolved spectroscopy and finite-difference time-domain simulations show robust emission for displacements up to $50$ nm (Stokes parameter $\lvert S \rvert < 0.05$ ). The devices display stable FSS and polarization imbalance below $5 \, \%$ , confirming precise, reproducible alignment and potential for high fidelity devices. This scalable approach enables deterministic integration of high-performance QDs with photonic cavities, advancing practical quantum light sources for quantum information technologies.

cond-mat.mtrl-sci

Free-standing circular Bragg gratings enabling efficient GaAs quantum dot entangled photon pair sources

Deterministic and bright quantum light sources based on scalable semiconductor technologies are a crucial building block for future quantum communication networks. While circular Bragg gratings (CBGs) are highly effective for extracting light from solid-state quantum emitters, conventional architectures rely on complex multi-layer processing or flip-chip bonding, which introduce detrimental strain and limit scalability. Here, we present a fabrication-minimal approach to realize monolithic, free-standing CBG cavities with deterministically positioned single GaAs quantum dots (QDs). By utilizing aspect-ratio-dependent etching (ARDE) in a single-step top-down process, we achieve the necessary vertical structural asymmetry for directional emission without requiring bottom reflectors. Finite-difference time-domain (FDTD) simulations validate this geometry, predicting free-space extraction efficiencies up to $68 \, \%$ and coupling efficiencies of $40 \, \%$ into a lensed single-mode fiber ($\text{NA} = 0.6$). Experimentally, the deterministically coupled QD-CBG devices yield a photoluminescence intensity enhancement of up to $\times 700$ compared to unprocessed planar QDs, reaching integrated count rates of $45 \, MHz$. Furthermore, the suspended membrane architecture effectively relaxes residual strain, significantly reducing the average exciton fine-structure splitting from $7.3 \, \mu eV$ in planar QDs to $1.3 \, \mu eV$ in the CBGs. Interferometric measurements confirm that the fabrication process preserves the optical quality of the emitters, with average coherence times of $70 \, ps$. By bridging optimized FDTD design with precise nanofabrication and robust optical performance, these results establish free-standing GaAs CBGs as a highly scalable platform for bright and coherent entangled photon pair sources.

physics.optics

Compact system development of efficient quantum-entangled photon sources towards deployable and industrial devices

Entangled photon pair sources are a key enabling technology for quantum communication and networking, yet their deployment beyond laboratory environments is hindered by system-level complexity, limited operational stability, and insufficient industry compatibility. Here, we demonstrate a rack-based, mobile quantum light source architecture based on a semiconductor quantum dot emitter that directly addresses these challenges through modular system integration and automated operation. The source generates polarization-entangled photon pairs with an entanglement negativity 2n of up to $0.98(1)$, confirming near-maximal entanglement quality. In continuous, hands-off operation over a six-hour time window, the system achieves an average single-photon emission rate of $697(8)$ kHz and a maximum rate of $740(7)$ kHz, while maintaining 2n-value of more than $95$ $\%$. These results are enabled by the integration of optical excitation, collection, cryogenic operation, and control electronics within a standardized rack footprint, together with automated monitoring. By demonstrating simultaneously high entanglement quality, sustained brightness, and long-term operational stability in an industry-aligned system architecture, this work advances semiconductor quantum dot sources toward deployable entangled photon sources for applied quantum photonics.

quant-ph

MAD: Motion Appearance Decoupling for efficient Driving World Models

Recent video diffusion models generate photorealistic, temporally coherent videos, yet they fall short as reliable world models for autonomous driving, where structured motion and physically consistent interactions are essential. Adapting these generalist video models to driving domains has shown promise but typically requires massive domain-specific data and costly fine-tuning. We propose an efficient adaptation framework that converts generalist video diffusion models into controllable driving world models with minimal supervision. The key idea is to decouple motion learning from appearance synthesis. First, the model is adapted to predict structured motion in a simplified form: videos of skeletonized agents and scene elements, focusing learning on physical and social plausibility. Then, the same backbone is reused to synthesize realistic RGB videos conditioned on these motion sequences, effectively "dressing" the motion with texture and lighting. This two-stage process mirrors a reasoning-rendering paradigm: first infer dynamics, then render appearance. Our experiments show this decoupled approach is exceptionally efficient: adapting SVD, we match prior SOTA models with less than 6% of their compute. Scaling to LTX, our MAD-LTX model outperforms all open-source competitors, and supports a comprehensive suite of text, ego, and object controls. Project page: https://vita-epfl.github.io/MAD-World-Model/

cs.CV

Temperature-dependent refractive index of AlGaAs for quantum-photonic devices near the bandgap

We present an experimental method to determine the refractive index of $Al_{x}Ga_{1-x}As$ (x = 0.0 - 0.5) from 300 K to 4 K across the 500 - 1100 nm wavelength range. The values are extracted from spectroscopically observed microcavity resonances in thin $Al_{x}Ga_{1-x}As$ membranes embedded between fully and partially reflective gold mirrors. Refined Varshni and Paessler models are used to describe temperature-dependent bandgap shifts and material composition. By tracking resonance shifts and benchmarking against finite-difference time-domain simulations, we derive the dispersive optical response with high precision. This yields a quantitatively improved analytical expression for the refractive index of $Al_{x}Ga_{1-x}As$ matching the experimental results with a coefficient of determination as high as $R^2=0.993$, enabling accurate modeling near the band edge at cryogenic temperatures. The method is straightforward and broadly applicable to other semiconductor systems, offering a valuable tool for the design of micro photonic devices such as quantum light sources.

physics.optics

Rethinking Visual Intelligence: Insights from Video Pretraining

Large language models (LLMs) have demonstrated that large-scale pretraining enables systems to adapt rapidly to new problems with little supervision in the language domain. This success, however, has not translated as effectively to the visual domain, where models, including LLMs, continue to struggle with compositional understanding, sample efficiency, and general-purpose problem-solving. We investigate Video Diffusion Models (VDMs) as a promising direction for bridging this gap. Pretraining on spatiotemporal data endows these models with strong inductive biases for structure and dynamics, which we hypothesize can support broad task adaptability. To test this, we design a controlled evaluation in which both a pretrained LLM and a pretrained VDM are equipped with lightweight adapters and presented with tasks in their natural modalities. Across benchmarks including ARC-AGI, ConceptARC, visual games, route planning, and cellular automata, VDMs demonstrate higher data efficiency than their language counterparts. Taken together, our results indicate that video pretraining offers inductive biases that support progress toward visual foundation models.

cs.CV

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models

Video Diffusion Models (VDMs) have emerged as powerful generative tools, capable of synthesizing high-quality spatiotemporal content. Yet, their potential goes far beyond mere video generation. We argue that the training dynamics of VDMs, driven by the need to model coherent sequences, naturally pushes them to internalize structured representations and an implicit understanding of the visual world. To probe the extent of this internal knowledge, we introduce a few-shot fine-tuning framework that repurposes VDMs for new tasks using only a handful of examples. Our method transforms each task into a visual transition, enabling the training of LoRA weights on short input-output sequences without altering the generative interface of a frozen VDM. Despite minimal supervision, the model exhibits strong generalization across diverse tasks, from low-level vision (for example, segmentation and pose estimation) to high-level reasoning (for example, on ARC-AGI). These results reframe VDMs as more than generative engines. They are adaptable visual learners with the potential to serve as the backbone for future foundation models in vision.

cs.CV

Bright quantum dot light sources using monolithic microlenses on gold back-reflectors

We present the fabrication process of bright $GaAs$ quantum dot (QD) photon sources by non-deterministic embedding into broadband monolithic $Al_{0.15}Ga_{0.85}As$ microlens arrays on gold-coated substrates. Arrays of cylindrical photoresist templates, with diameters ranging from $2$ $\mu m$ to $5$ $\mu m$, are thermally reflowed and subsequently transferred into the $Al_{0.15}Ga_{0.85}As$ thin-film semiconductor heterostructure with embedded quantum dots through an optimized anisotropic and three-dimensional shape-preserving reactive ion etching process. This methodology facilitated the fabrication of large-scale ($2$ $mm$ $\times$ $4$ $mm$) and densely packed arrays of uniformly shaped microlenses ($\sim$ $40 \times 10^3$ $mm^{-1}$), with the brightest emissions from QDs embedded in microlenses exhibiting lateral diameters and heights of $2.7$ $\mu m$ and $1.35$ $\mu m$, respectively. Finite-difference time-domain simulations of both idealized and fabricated lens shapes provide a comprehensive three-dimensional analysis of the device performance and optimization potentials such as anti-reflection coatings. It is found that free-space extraction (fiber-coupled) efficiencies of up to $62$ $\%$ ($37$ $\%$) are achievable for hemispherical QD-microlenses on gold-coated substrates. A statistical model for the fabrication yield of QD-microlenses is developed and experimentally corroborated by photoluminescence spectroscopy of fabricated microlens arrays. This analysis exhibited a free-space intensity enhancement by factors of up to $\times 200$ in approximately $1$ out of $200$ microlenses, showing good agreement to the theoretical expectations. This scalable fabrication strategy underscores the potential of these compact, high-efficiency sources offering new prospects for applications of these devices in future large-scale quantum networks.

cond-mat.mes-hall

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth outputs for richer spatial understanding. We introduce autoregressive noise schedules to enable stable long-horizon generations. Our dataset is comprised of 4000+ hours of multimodal data across domains like autonomous driving, egocentric human activities, and drone flights. Pseudo-labels are used to get depth maps, ego-trajectories, and human poses. We use a comprehensive evaluation framework, including a new Control of Object Manipulation (COM) metric, to assess controllability. Experiments show GEM excels at generating diverse, controllable scenarios and temporal consistency over long generations. Code, models, and datasets are fully open-sourced.

cs.CV

A Multi-Loss Strategy for Vehicle Trajectory Prediction: Combining Off-Road, Diversity, and Directional Consistency Losses

Trajectory prediction is essential for the safety and efficiency of planning in autonomous vehicles. However, current models often fail to fully capture complex traffic rules and the complete range of potential vehicle movements. Addressing these limitations, this study introduces three novel loss functions: Offroad Loss, Direction Consistency Error, and Diversity Loss. These functions are designed to keep predicted paths within driving area boundaries, aligned with traffic directions, and cover a wider variety of plausible driving scenarios. As all prediction modes should adhere to road rules and conditions, this work overcomes the shortcomings of traditional "winner takes all" training methods by applying the loss functions to all prediction modes. These loss functions not only improve model training but can also serve as metrics for evaluating the realism and diversity of trajectory predictions. Extensive validation on the nuScenes and Argoverse 2 datasets with leading baseline models demonstrates that our approach not only maintains accuracy but significantly improves safety and robustness, reducing offroad errors on average by 47% on original and by 37% on attacked scenes. This work sets a new benchmark for trajectory prediction in autonomous driving, offering substantial improvements in navigating complex environments. Our code is available at https://github.com/vita-epfl/stay-on-track .

cs.CV

Sim-to-Real Causal Transfer: A Metric Learning Approach to Causally-Aware Interaction Representations

Modeling spatial-temporal interactions among neighboring agents is at the heart of multi-agent problems such as motion forecasting and crowd navigation. Despite notable progress, it remains unclear to which extent modern representations can capture the causal relationships behind agent interactions. In this work, we take an in-depth look at the causal awareness of these representations, from computational formalism to real-world practice. First, we cast doubt on the notion of non-causal robustness studied in the recent CausalAgents benchmark. We show that recent representations are already partially resilient to perturbations of non-causal agents, and yet modeling indirect causal effects involving mediator agents remains challenging. To address this challenge, we introduce a metric learning approach that regularizes latent representations with causal annotations. Our controlled experiments show that this approach not only leads to higher degrees of causal awareness but also yields stronger out-of-distribution robustness. To further operationalize it in practice, we propose a sim-to-real causal transfer method via cross-domain multi-task learning. Experiments on pedestrian datasets show that our method can substantially boost generalization, even in the absence of real-world causal annotations. We hope our work provides a new perspective on the challenges and pathways towards causally-aware representations of multi-agent interactions. Our code is available at https://github.com/vita-epfl/CausalSim2Real.

cs.LG

Vehicle trajectory prediction works, but not everywhere

Vehicle trajectory prediction is nowadays a fundamental pillar of self-driving cars. Both the industry and research communities have acknowledged the need for such a pillar by providing public benchmarks. While state-of-the-art methods are impressive, i.e., they have no off-road prediction, their generalization to cities outside of the benchmark remains unexplored. In this work, we show that those methods do not generalize to new scenes. We present a method that automatically generates realistic scenes causing state-of-the-art models to go off-road. We frame the problem through the lens of adversarial scene generation. The method is a simple yet effective generative model based on atomic scene generation functions along with physical constraints. Our experiments show that more than 60% of existing scenes from the current benchmarks can be modified in a way to make prediction methods fail (i.e., predicting off-road). We further show that the generated scenes (i) are realistic since they do exist in the real world, and (ii) can be used to make existing models more robust, yielding 30-40 reductions in the off-road rate. The code is available online: https://s-attack.github.io/.

cs.CV