SearcharxivSearch

arXiv subjects

Dmitry Kalashnikov

Publications and source records attributed to Dmitry Kalashnikov.

At least 19 recordsLinked to original sources

Multimode squeezed light generation and characterization

Nowadays, the realization of quantum computations and communications based on continuous variables has attracted a significant attention due to a substantial expansion of the system dimensionality. The main progress in this area is attributed to the implementation of multimode systems based on squeezed states of light. One of the simplest ways to generate such states relies upon their producing in a single-pass optical parametric amplifier (OPA) using ultrafast pumping. However, for homodyne detection of such multimode states, the profile of the local oscillator (LO) must perfectly match the profile of the measured mode. Usually, this is not the case; therefore a proper treatment of multimode squeezing is required. In this work, we study both theoretically and experimentally the multimode squeezed light generated in type-0 and type-II OPA. We characterize such sources and investigate the degree of squeezing in dependence on the LO spectral profile, employing a pulse shaping technique. The theoretical analysis is performed using the Schmidt-mode theory. This work might have a significant impact on the realization of multimode quantum protocols

quant-ph

Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the Gemini Robotics model family: Gemini Robotics 1.5, a multi-embodiment Vision-Language-Action (VLA) model, and Gemini Robotics-ER 1.5, a state-of-the-art Embodied Reasoning (ER) model. We are bringing together three major innovations. First, Gemini Robotics 1.5 features a novel architecture and a Motion Transfer (MT) mechanism, which enables it to learn from heterogeneous, multi-embodiment robot data and makes the VLA more general. Second, Gemini Robotics 1.5 interleaves actions with a multi-level internal reasoning process in natural language. This enables the robot to "think before acting" and notably improves its ability to decompose and execute complex, multi-step tasks, and also makes the robot's behavior more interpretable to the user. Third, Gemini Robotics-ER 1.5 establishes a new state-of-the-art for embodied reasoning, i.e., for reasoning capabilities that are critical for robots, such as visual and spatial understanding, task planning, and progress estimation. Together, this family of models takes us a step towards an era of physical agents-enabling robots to perceive, think and then act so they can solve complex multi-step tasks.

cs.RO

Can AI Perceive Physical Danger and Intervene?

When AI interacts with the physical world -- as a robot or an assistive agent -- new safety challenges emerge beyond those of purely ``digital AI". In such interactions, the potential for physical harm is direct and immediate. How well do state-of-the-art foundation models understand common-sense facts about physical safety, e.g. that a box may be too heavy to lift, or that a hot cup of coffee should not be handed to a child? In this paper, our contributions are three-fold: first, we develop a highly scalable approach to continuous physical safety benchmarking of Embodied AI systems, grounded in real-world injury narratives and operational safety constraints. To probe multi-modal safety understanding, we turn these narratives and constraints into photorealistic images and videos capturing transitions from safe to unsafe states, using advanced generative models. Secondly, we comprehensively analyze the ability of major foundation models to perceive risks, reason about safety, and trigger interventions; this yields multi-faceted insights into their deployment readiness for safety-critical agentic applications. Finally, we develop a post-training paradigm to teach models to explicitly reason about embodiment-specific safety constraints provided through system instructions. The resulting models generate thinking traces that make safety reasoning interpretable and transparent, achieving state of the art performance in constraint satisfaction evaluations. The benchmark is released at https://asimov-benchmark.github.io/v2

cs.AI

Searches for new light particles at the Troitsk Meson Factory (TiMoFey)

The project of a new accelerator complex at the Institute for Nuclear Research of RAS in Troitsk has recently been included in the Russian National Program ``Fundamental Properties of Matter". It will sustain a proton beam with a current of 300 (100) $\mu$A and a proton kinetic energy of $T_p=423\,(1300)$ MeV at the first (second) stage of operation. The complex is multidisciplinary, and here we investigate its prospects in exploring new physics with light, feebly interacting particles. We find that TiMoFey can access new regions of parameter space of models with light axion-like particles and models with hidden photons, provided by a generic multipurpose detector installed downstream the proton beam dump. The signature to be exploited is the decay of a new particle into a pair of known particles inside the detector. Likewise, TiMoFey can probe previously unreachable ranges of parameters of models with millicharged particles obtained in measurements with detectors recognizing energy deposits associated with elastic scattering of new particles, passing through the detector volume. The latter detector may be useful for dark matter searches, as well as for studies of neutrino physics suggested at the facility in the Program framework.

hep-ph

Gemini Robotics: Bringing AI into the Physical World

Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as robots remains a significant challenge. This report introduces a new family of AI models purposefully designed for robotics and built upon the foundation of Gemini 2.0. We present Gemini Robotics, an advanced Vision-Language-Action (VLA) generalist model capable of directly controlling robots. Gemini Robotics executes smooth and reactive movements to tackle a wide range of complex manipulation tasks while also being robust to variations in object types and positions, handling unseen environments as well as following diverse, open vocabulary instructions. We show that with additional fine-tuning, Gemini Robotics can be specialized to new capabilities including solving long-horizon, highly dexterous tasks, learning new short-horizon tasks from as few as 100 demonstrations and adapting to completely novel robot embodiments. This is made possible because Gemini Robotics builds on top of the Gemini Robotics-ER model, the second model we introduce in this work. Gemini Robotics-ER (Embodied Reasoning) extends Gemini's multimodal reasoning capabilities into the physical world, with enhanced spatial and temporal understanding. This enables capabilities relevant to robotics including object detection, pointing, trajectory and grasp prediction, as well as multi-view correspondence and 3D bounding box predictions. We show how this novel combination can support a variety of robotics applications. We also discuss and address important safety considerations related to this new class of robotics foundation models. The Gemini Robotics family marks a substantial step towards developing general-purpose robots that realizes AI's potential in the physical world.

cs.RO

Generating Robot Constitutions & Benchmarks for Semantic Safety

Until recently, robotics safety research was predominantly about collision avoidance and hazard reduction in the immediate vicinity of a robot. Since the advent of large vision and language models (VLMs), robots are now also capable of higher-level semantic scene understanding and natural language interactions with humans. Despite their known vulnerabilities (e.g. hallucinations or jail-breaking), VLMs are being handed control of robots capable of physical contact with the real world. This can lead to dangerous behaviors, making semantic safety for robots a matter of immediate concern. Our contributions in this paper are two fold: first, to address these emerging risks, we release the ASIMOV Benchmark, a large-scale and comprehensive collection of datasets for evaluating and improving semantic safety of foundation models serving as robot brains. Our data generation recipe is highly scalable: by leveraging text and image generation techniques, we generate undesirable situations from real-world visual scenes and human injury reports from hospitals. Secondly, we develop a framework to automatically generate robot constitutions from real-world data to steer a robot's behavior using Constitutional AI mechanisms. We propose a novel auto-amending process that is able to introduce nuances in written rules of behavior; this can lead to increased alignment with human preferences on behavior desirability and safety. We explore trade-offs between generality and specificity across a diverse set of constitutions of different lengths, and demonstrate that a robot is able to effectively reject unconstitutional actions. We measure a top alignment rate of 84.3% on the ASIMOV Benchmark using generated constitutions, outperforming no-constitution baselines and human-written constitutions. Data is available at asimov-benchmark.github.io

cs.RO

Playing with lepton asymmetry at the resonant production of sterile neutrino dark matter

We examine the sterile neutrino dark matter production in the primordial plasma with lepton asymmetry unequally distributed over different neutrino flavors. We argue that with the specific flavor fractions, one can mitigate limits from the Big Bang Nucleosynthesis on the sterile-active neutrino mixing angle and sterile neutrino mass. It happens due to cancellation of the neutrino flavor asymmetries in active neutrino oscillations, which is more efficient in the case of inverse hierarchy of active neutrino masses and does not depend on the value of CP-phase. This finding opens a window of lower sterile-active mixing angles. Likewise, we show that, with lepton asymmetry disappearing from the plasma at certain intermediate stages of the sterile neutrino production, the spectrum of produced neutrinos becomes much colder, which weakens the limits on the model parameter space from observations of cosmic small-scale structures (Ly-$\alpha$ forest, galaxy counts, etc.). This finding reopens the region of lighter sterile neutrinos. The new region may be explored with the next generation of X-ray telescopes searching for the inherent peak signature provided by the dark matter sterile neutrino radiative decays in the Galaxy.

hep-ph

Predictive Red Teaming: Breaking Policies Without Breaking Robots

Visuomotor policies trained via imitation learning are capable of performing challenging manipulation tasks, but are often extremely brittle to lighting, visual distractors, and object locations. These vulnerabilities can depend unpredictably on the specifics of training, and are challenging to expose without time-consuming and expensive hardware evaluations. We propose the problem of predictive red teaming: discovering vulnerabilities of a policy with respect to environmental factors, and predicting the corresponding performance degradation without hardware evaluations in off-nominal scenarios. In order to achieve this, we develop RoboART: an automated red teaming (ART) pipeline that (1) modifies nominal observations using generative image editing to vary different environmental factors, and (2) predicts performance under each variation using a policy-specific anomaly detector executed on edited observations. Experiments across 500+ hardware trials in twelve off-nominal conditions for visuomotor diffusion policies demonstrate that RoboART predicts performance degradation with high accuracy (less than 0.19 average difference between predicted and real success rates). We also demonstrate how predictive red teaming enables targeted data collection: fine-tuning with data collected under conditions predicted to be adverse boosts baseline performance by 2-7x.

cs.RO

Learning the RoPEs: Better 2D and 3D Position Encodings with STRING

We introduce STRING: Separable Translationally Invariant Position Encodings. STRING extends Rotary Position Encodings, a recently proposed and widely used algorithm in large language models, via a unifying theoretical framework. Importantly, STRING still provides exact translation invariance, including token coordinates of arbitrary dimensionality, whilst maintaining a low computational footprint. These properties are especially important in robotics, where efficient 3D token representation is key. We integrate STRING into Vision Transformers with RGB(-D) inputs (color plus optional depth), showing substantial gains, e.g. in open-vocabulary object detection and for robotics controllers. We complement our experiments with a rigorous mathematical analysis, proving the universality of our methods.

cs.LG

STEER: Flexible Robotic Manipulation via Dense Language Grounding

The complexity of the real world demands robotic systems that can intelligently adapt to unseen situations. We present STEER, a robot learning framework that bridges high-level, commonsense reasoning with precise, flexible low-level control. Our approach translates complex situational awareness into actionable low-level behavior through training language-grounded policies with dense annotation. By structuring policy training around fundamental, modular manipulation skills expressed in natural language, STEER exposes an expressive interface for humans or Vision-Language Models (VLMs) to intelligently orchestrate the robot's behavior by reasoning about the task and context. Our experiments demonstrate the skills learned via STEER can be combined to synthesize novel behaviors to adapt to new situations or perform completely new tasks without additional data collection or training.

cs.RO

Dielectric Fano Nanoantennas for Enabling Sub-Nanosecond Lifetimes in NV-based Single Photon Emitters

Solid-state quantum emitters are essential sources of single photons, and enhancing their emission rates is of paramount importance for applications in quantum communications, computing, and metrology. One approach is to couple quantum emitters with resonant photonic nanostructures, where the emission rate is enhanced due to the Purcell effect. Dielectric nanoantennas are promising as they provide strong emission enhancement compared to plasmonic ones, which suffer from high Ohmic loss. Here, we designed and fabricated a dielectric Fano resonator based on a pair of silicon (Si) ellipses and a disk, which supports the mode hybridization between quasi-bound-states-in-the-continuum (quasi-BIC) and Mie resonance. We demonstrated the performance of the developed resonant system by interfacing it with single photon emitters (SPEs) based on nitrogen-vacancy (NV-) centers in nanodiamonds (NDs). We observed that the interfaced emitters have a Purcell enhancement factor of ~10, with sub-ns emission lifetime and a polarization contrast of 9. Our results indicate a promising method for developing efficient and compact single-photon sources for integrated quantum photonics applications.

physics.optics

NICA prospects in searches for light exotics from hidden sectors: the cases of hidden photons and axion-like particles

We present first estimates of NICA sensitivity to Standard Model extensions with light hypothetical particles singlet under the known gauge transformations. Our analysis reveals that NICA can explore new regions in the parameter spaces of models with a hidden vector and models with an axion-like particle of masses about 30-500\,MeV. Some of these regions seem unreachable by other ongoing and approved future projects. NICA has good prospects in discovery ($5\sigma$) of the new physics after 1 year of data taking.

hep-ph

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train generalist X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. More details can be found on the project website https://robotics-transformer-x.github.io.

cs.RO

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions and enjoy the benefits of large-scale pretraining on language and vision-language data from the web. To this end, we propose to co-fine-tune state-of-the-art vision-language models on both robotic trajectory data and Internet-scale vision-language tasks, such as visual question answering. In contrast to other approaches, we propose a simple, general recipe to achieve this goal: in order to fit both natural language responses and robotic actions into the same format, we express the actions as text tokens and incorporate them directly into the training set of the model in the same way as natural language tokens. We refer to such category of models as vision-language-action models (VLA) and instantiate an example of such a model, which we call RT-2. Our extensive evaluation (6k evaluation trials) shows that our approach leads to performant robotic policies and enables RT-2 to obtain a range of emergent capabilities from Internet-scale training. This includes significantly improved generalization to novel objects, the ability to interpret commands not present in the robot training data (such as placing an object onto a particular number or icon), and the ability to perform rudimentary reasoning in response to user commands (such as picking up the smallest or largest object, or the one closest to another object). We further show that incorporating chain of thought reasoning allows RT-2 to perform multi-stage semantic reasoning, for example figuring out which object to pick up for use as an improvised hammer (a rock), or which type of drink is best suited for someone who is tired (an energy drink).

cs.RO

RT-1: Robotics Transformer for Real-World Control at Scale

By transferring knowledge from large, diverse, task-agnostic datasets, modern machine learning models can solve specific downstream tasks either zero-shot or with small task-specific datasets to a high level of performance. While this capability has been demonstrated in other fields such as computer vision, natural language processing or speech recognition, it remains to be shown in robotics, where the generalization capabilities of the models are particularly critical due to the difficulty of collecting real-world robotic data. We argue that one of the keys to the success of such general robotic models lies with open-ended task-agnostic training, combined with high-capacity architectures that can absorb all of the diverse, robotic data. In this paper, we present a model class, dubbed Robotics Transformer, that exhibits promising scalable model properties. We verify our conclusions in a study of different model classes and their ability to generalize as a function of the data size, model size, and data diversity based on a large-scale data collection on real robots performing real-world tasks. The project's website and videos can be found at robotics-transformer1.github.io

cs.RO

Probing light exotics from a hidden sector at $c$-$\tau$ factories with polarized electron beams

Future $c$-$\tau$ factories are natural places to study extensions of the Standard Model of particle physics (SM) with new long-lived feebly interacting particles light enough to be produced in electron-positron collisions. We investigate prospects of these machines in exploring such extensions emphasizing the role of polarized beams in getting rid of the SM irreducible background for the missing energy signature. We illustrate this on example of $c$-$\tau$ project in Novosibirsk, where the electron beam is designed to be polarized to achieve much higher sensitivity to hadronic resonances and $\tau$-leptons. We investigate models with hidden photons, with millicharged particles (fermions and scalars), with $Z'$ bosons and with axion-like particles. We find that the electron beam polarization of 80\% significantly improves the chances to observe the signal, especially with large statistics. We outline the regions of the model parameter space which can be reached at this factory in one year and in ten years of operation according with the scientific schedule of tuning the energy of colliding beams.

hep-ph

Learning Model Predictive Controllers with Real-Time Attention for Real-World Navigation

Despite decades of research, existing navigation systems still face real-world challenges when deployed in the wild, e.g., in cluttered home environments or in human-occupied public spaces. To address this, we present a new class of implicit control policies combining the benefits of imitation learning with the robust handling of system constraints from Model Predictive Control (MPC). Our approach, called Performer-MPC, uses a learned cost function parameterized by vision context embeddings provided by Performers -- a low-rank implicit-attention Transformer. We jointly train the cost function and construct the controller relying on it, effectively solving end-to-end the corresponding bi-level optimization problem. We show that the resulting policy improves standard MPC performance by leveraging a few expert demonstrations of the desired navigation behavior in different challenging real-world scenarios. Compared with a standard MPC policy, Performer-MPC achieves >40% better goal reached in cluttered environments and >65% better on social metrics when navigating around humans.

cs.RO

On direct observation of millicharged particles at $c$-$\tau$ factories and other $e^+e^-$-colliders

Hypothetical particles with tiny electric charges (millicharged particles or MCPs) can be produced in electron-positron annihilation if kinematically allowed. Typical searches for them at $e^+e^-$ colliders exploit a signature of a single photon with missing energy carried away by the undetected MCP pair. We put forward an idea to look alternatively for MCP energy deposits inside a tracker, which is a direct observation. The new signature is relevant for non-relativistic MCPs, and we illustrate its power on the example of the $c$-$\tau$ factory, where we argued that the corresponding searches may be background-free. We find that it can probe the MCP charge down to $3\times10^{-3}$ of the electron charge for the MCP masses in ${\cal O}(5)$ MeV vicinity of each energy beam value where the factory will collect a luminosity of 100 fb$^{-1}$ in one year. This mass region is unreachable with the searches for missing energy and single photon.

hep-ph