SearcharxivSearch

arXiv subjects

Fei-Fei Li

Publications and source records attributed to Fei-Fei Li.

12 recordsLinked to original sources

One Demonstration, Many Objects: Generalizing Manipulation via Local Contact Geometry

Dexterous manipulation with multi-fingered robot hands promises human-level dexterity, but collecting large-scale dexterous robot hand data remains difficult. Learning from human demonstrations has emerged as a scalable alternative to robot teleoperation, providing strong priors on object interaction and contact strategies. Recent sim-to-real RL methods incorporate such priors, but often (i) omit rewards that explicitly incentivize precise contact, yielding weak real-world performance, and/or (ii) generalize poorly to unseen object instances. We propose DemoMimic (Dexterous Motion Mimic), a policy that manipulates objects by focusing on their geometry local to the contact points. Its contact-centric rewards encourage precise contact and improve sim-to-real consistency, yielding a single real-world policy that transfers across objects of varying shape, scale, mass, and friction wherever local contact structure is preserved. Real-world ablations show that DemoMimic achieves 71% success across 16 objects, four tasks, and two robot-hand embodiments, with the smallest sim-to-real drop compared to baselines.

cs.RO

Antichiral hinge states in a higher-order photonic nodal ring semimetal

Antichiral states propagate in the same direction on opposite boundaries, defying the conventional constraint that boundary modes must cancel net chirality. Previously found antichiral states have been limited to first-order topological semimetals. However, antichiral hinge states, the antichiral counterpart of recently discovered higher-order chiral hinge states, remain elusive. Here, we report the observation of antichiral hinge states in a higher-order nodal-ring semimetal made of a three-dimensional gyromagnetic photonic crystal. Near-field scanning measurements reveal a pair of hinge states at two parallel one-dimensional boundaries propagating unidirectionally along the same direction, with robust transport against metallic scatterers. Their spatial positions can be reconfigured by adding or removing photonic layers. The additionally observed antichiral and drumhead surface states manifest a hierarchy of first- and second-order topological boundary states within a single photonic system. Our work extends antichiral states to higher-order topological semimetals and has potential applications in robust and reconfigurable photonic routing.

physics.optics

Representations of shifted twisted quantum affine algebras

In this paper, we introduce and study shifted twisted quantum affine algebras which provide a twisted counterpart of the theory of shifted quantum affine algebras. The shifted twisted quantum affine algebra $\U_q^{\mu_+,\mu_-}(\hgs)$ is obtained from the Drinfeld current presentation of twisted quantum loop algebras by shifting the Cartan--Drinfeld currents $\phi_i^\pm(z)$ according to a coweight pair $(\mu_+,\mu_-)$. We prove that it admits a triangular decomposition and that, up to isomorphism, they depend only on the total shift $\mu=\mu_+ + \mu_-$. For each shift $\mu$, we define a category $\mathcal O_\mu$ of representations of $\U_q^\mu(\hgs) = \U_q^{0,\mu}(\hgs)$ and prove a rationality theorem for the Cartan currents: on every weight space, the two currents $\phi_i^+(z)$ and $\phi_i^-(z)$ are expansions of the same rational operator-valued function, whose degree is prescribed by $\alpha_i(\mu)$. As a consequence, we classify the simple objects of $\mathcal O_\mu$ by rational $\ell$-weights of the corresponding degrees. We then construct a deformed Drinfeld coproduct and use it to define a fusion product on the direct sum $\mathcal{O}^{sh}$ of the categories $\mathcal O_\mu$. This fusion product is compatible with $q$-characters. We also classify finite-dimensional simple modules in $\mathcal{O}^{sh}$ in terms of dominant rational $\ell$-weights, with a separate treatment of type $A_{2n}^{(2)}$. Finally, we construct restriction representations relating representations of twisted quantum affine Borel algebras to representations of shifted twisted quantum affine algebras, and establish a $q$-characters formula for simple finite-dimensional representations of shifted twisted quantum affine algebras in terms of the $q$-characters of the corresponding simple representations of the twisted quantum affine Borel algebra $\U_q(\bs)$.

math.QA

MIRAGE: The Illusion of Visual Understanding

Multimodal AI systems have achieved remarkable performance across a broad range of real-world tasks, yet the mechanisms underlying visual-language reasoning remain surprisingly poorly understood. We report three findings that challenge prevailing assumptions about how these systems process and integrate visual information. First, Frontier models readily generate detailed image descriptions and elaborate reasoning traces, including pathology-biased clinical findings, for images never provided; we term this phenomenon mirage reasoning. Second, without any image input, models also attain strikingly high scores across general and medical multimodal benchmarks, bringing into question their utility and design. In the most extreme case, our model achieved the top rank on a standard chest X-ray question-answering benchmark without access to any images. Third, when models were explicitly instructed to guess answers without image access, rather than being implicitly prompted to assume images were present, performance declined markedly. Explicit guessing appears to engage a more conservative response regime, in contrast to the mirage regime in which models behave as though images have been provided. These findings expose fundamental vulnerabilities in how visual-language models reason and are evaluated, pointing to an urgent need for private benchmarks that eliminate textual cues enabling non-visual inference, particularly in medical contexts where miscalibrated AI carries the greatest consequence. We introduce B-Clean as a principled solution for fair, vision-grounded evaluation of multimodal AI systems.

cs.AI

CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation

"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaP-X, an open-access framework for systematically studying Code-as-Policy agents in robot manipulation. At its core is CaP-Gym, an interactive environment in which agents control robots by synthesizing and executing programs that compose perception and control primitives. Building on this foundation, CaP-Bench evaluates frontier language and vision-language models across varying levels of abstraction, interaction, and perceptual grounding. Across 12 models, CaP-Bench reveals a consistent trend: performance improves with human-crafted abstractions but degrades as these priors are removed, exposing a dependence on designer scaffolding. At the same time, we observe that this gap can be mitigated through scaling agentic test-time computation--through multi-turn interaction, structured execution feedback, visual differencing, automatic skill synthesis, and ensembled reasoning--substantially improves robustness even when agents operate over low-level primitives. These findings allow us to derive CaP-Agent0, a training-free framework that recovers human-level reliability on several manipulation tasks in simulation and on real embodiments. We further introduce CaP-RL, showing reinforcement learning with verifiable rewards improves success rates and transfers from sim2real with minimal gap. Together, CaP-X provides a principled, open-access platform for advancing embodied coding agents.

cs.RO

Mineral Detection of Cosmic-Ray Boosted Dark Matter

We present the first dedicated analysis of cosmic-ray boosted dark matter (CRDM) in paleo detectors. Owing to their large kinetic energies, CRDM particles generate nuclear-recoil tracks that extend to substantially larger lengths than those produced by dominant backgrounds from neutrinos and intrinsic radioactivity. Combined with the ultra-large effective geological exposure of $\mathcal{O}(10^{5})~\mathrm{t\,yr}$, paleo detectors provide a uniquely sensitive probe of sub-GeV DM. Considering both constant and vector-mediator interactions, we find that paleo detectors improve the sensitivity to the DM--proton scattering cross section by one to two orders of magnitude compared with the latest XENONnT limits.

hep-ph

Observation of fractional topological numbers at photonic edges and corners

Topological phases of matter are featured with exotic edge states. However, the fractional topological numbers at edges, though predicted long ago by Jackiw and Rebbi, remain elusive in topological photonic systems. Here, we report on the observation of fractional topological numbers at the topological edges and corners in one- and two-dimensional photonic crystals. The fractional topological numbers are determined via the measurements of the photonic local density-of-states. In one-dimensional photonic crystals, we witness a rapid change of the fractional topological number at the edges rising from 0 to 1/2 when the photonic band gap experiences a topological transition, confirming the well-known prediction of Jackiw and Rebbi. In two-dimensional systems, we discover that the fractional topological number in the corner region varies from 0 to 1/2 and 1/4 in different photonic band gap phases. Our study paves the way toward topological manipulation of fractional quantum numbers in photonics.

physics.optics

Experimental discovery of bulk-disclination correspondence

Most natural and artificial materials have crystalline structures from which abundant topological phases emerge [1-6]. The bulk-edge correspondence, widely-adopted in experiments to determine the band topology from edge properties, however, becomes inadequate in discerning various topological crystalline phases [7-17], leading to great challenges in the experimental classification of the large family of topological crystalline materials [4-6]. Theories predict that disclinations, ubiquitous crystallographic defects, provide an effective probe of crystalline topology beyond edges [18-21], which, however, has not yet been confirmed in experiments. Here, we report the experimental discovery of the bulk-disclination correspondence which is manifested as the fractional spectral charge and robust bound states at the disclinations. The fractional disclination charge originates from the symmetry-protected bulk charge patterns---a fundamental property of many topological crystalline insulators (TCIs). Meanwhile, the robust bound states at disclinations emerge as a secondary, but directly observable property of TCIs. Using reconfigurable photonic crystals as photonic TCIs with higher-order topology, we observe those hallmark features via pump-probe and near-field detection measurements. Both the fractional charge and the localized states are demonstrated to emerge at the disclination in the TCI phase but vanish in the trivial phase. The experimental discovery of bulk-disclination correspondence unveils a novel fundamental phenomenon and a new paradigm for exploring topological materials.

cond-mat.mtrl-sci

DDRprog: A CLEVR Differentiable Dynamic Reasoning Programmer

We present a novel Dynamic Differentiable Reasoning (DDR) framework for jointly learning branching programs and the functions composing them; this resolves a significant nondifferentiability inhibiting recent dynamic architectures. We apply our framework to two settings in two highly compact and data efficient architectures: DDRprog for CLEVR Visual Question Answering and DDRstack for reverse Polish notation expression evaluation. DDRprog uses a recurrent controller to jointly predict and execute modular neural programs that directly correspond to the underlying question logic; it explicitly forks subprocesses to handle logical branching. By effectively leveraging additional structural supervision, we achieve a large improvement over previous approaches in subtask consistency and a small improvement in overall accuracy. We further demonstrate the benefits of structural supervision in the RPN setting: the inclusion of a stack assumption in DDRstack allows our approach to generalize to long expressions where an LSTM fails the task.

cs.CV

Topological light-trapping on a dislocation

Topology has been revealed to play a fundamental role in physics in the past decades. Topological insulators have unconventional gapless edge states where disorder-induced back-scattering is suppressed. In photonics, such edge states lead to unidirectional waveguides which are useful for integrated photonic chips. Cavity modes, another type of fundamental components in photonic chips, however, are not protected by band topology because of their lower dimensions. Here we demonstrate that concurrent wavevector-space and real-space topology, dubbed as the "dual-topology", can lead to light-trapping in lower-dimensions. The resultant photonic bound state emerges as a Jackiw-Rebbi soliton mode localized on a dislocation in a two-dimensional (2D) photonic crystal, as predicted theoretically and discovered experimentally. Such a strongly-confined 0D localized mode, which is solely due to the topological mechanism, is found to be robust against perturbations. Our study unveils a new mechanism for topological light-trapping in lower-dimensions, which is valuable for fundamental physics and a variety of applications in photonics.

physics.optics

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Despite progress in perceptual tasks such as image classification, computers still perform poorly on cognitive tasks such as image description and question answering. Cognition is core to tasks that involve not just recognizing, but reasoning about our visual world. However, models used to tackle the rich content in images for cognitive tasks are still being trained using the same datasets designed for perceptual tasks. To achieve success at cognitive tasks, models need to understand the interactions and relationships between objects in an image. When asked "What vehicle is the person riding?", computers will need to identify the objects in an image as well as the relationships riding(man, carriage) and pulling(horse, carriage) in order to answer correctly that "the person is riding a horse-drawn carriage". In this paper, we present the Visual Genome dataset to enable the modeling of such relationships. We collect dense annotations of objects, attributes, and relationships within each image to learn these models. Specifically, our dataset contains over 100K images where each image has an average of 21 objects, 18 attributes, and 18 pairwise relationships between objects. We canonicalize the objects, attributes, relationships, and noun phrases in region descriptions and questions answer pairs to WordNet synsets. Together, these annotations represent the densest and largest dataset of image descriptions, objects, attributes, relationships, and question answers.

cs.CV

Love Thy Neighbors: Image Annotation by Exploiting Image Metadata

Some images that are difficult to recognize on their own may become more clear in the context of a neighborhood of related images with similar social-network metadata. We build on this intuition to improve multilabel image annotation. Our model uses image metadata nonparametrically to generate neighborhoods of related images using Jaccard similarities, then uses a deep neural network to blend visual information from the image and its neighbors. Prior work typically models image metadata parametrically, in contrast, our nonparametric treatment allows our model to perform well even when the vocabulary of metadata changes between training and testing. We perform comprehensive experiments on the NUS-WIDE dataset, where we show that our model outperforms state-of-the-art methods for multilabel image annotation even when our model is forced to generalize to new types of metadata.

cs.CV