Searcharxiv⌕ Search

arXiv subjects

Eric Chen

Publications and source records attributed to Eric Chen.

At least 37 records · Page 2Linked to original sources

Learning-on-the-Drive: Self-supervised Adaptation of Visual Offroad Traversability Models

Autonomous offroad driving is essential for applications like emergency rescue, military operations, and agriculture. Despite progress, systems struggle with high-speed vehicles exceeding 10m/s due to the need for accurate long-range (> 50m) perception for safe navigation. Current approaches are limited by sensor constraints; LiDAR-based methods offer precise short-range data but are noisy beyond 30m, while visual models provide dense long-range measurements but falter with unseen scenarios. To overcome these issues, we introduce ALTER, a learning-on-the-drive perception framework that leverages both sensor types. ALTER uses a self-supervised visual model to learn and adapt from near-range LiDAR measurements, improving long-range prediction in new environments without manual labeling. It also includes a model selection module for better sensor failure response and adaptability to known environments. Testing in two real-world settings showed on average 43.4% better traversability prediction than LiDAR-only and 164% over non-adaptive state-of-the-art (SOTA) visual semantic methods after 45 seconds of online learning.

cs.RO↗

SmartChoices: Augmenting Software with Learned Implementations

In many software systems, heuristics are used to make decisions - such as cache eviction, task scheduling, and information presentation - that have a significant impact on overall system behavior. While machine learning may outperform these heuristics, replacing existing heuristics in a production system safely and reliably can be prohibitively costly. We present SmartChoices, a novel approach that reduces the cost to deploy production-ready ML solutions for contextual bandits problems. SmartChoices' interface cleanly separates problem formulation from implementation details: engineers describe their use case by defining datatypes for the context, arms, and feedback that are passed to SmartChoices APIs, while SmartChoices manages encoding & logging data and training, evaluating & deploying policies. Our implementation codifies best practices, is efficient enough for use in low-level applications, and provides valuable production features off the shelf via a shared library. Overall, SmartChoices enables non-experts to rapidly deploy production-ready ML solutions by eliminating many sources of technical debt common to ML systems. Engineers have independently used SmartChoices to improve a wide range of software including caches, batch processing workloads, and UI layouts, resulting in better latency, throughput, and click-through rates.

cs.SE↗

Grid Cell-Inspired Fragmentation and Recall for Efficient Map Building

Animals and robots navigate through environments by building and refining maps of space. These maps enable functions including navigation back to home, planning, search and foraging. Here, we use observations from neuroscience, specifically the observed fragmentation of grid cell map in compartmentalized spaces, to propose and apply the concept of Fragmentation-and-Recall (FARMap) in the mapping of large spaces. Agents solve the mapping problem by building local maps via a surprisal-based clustering of space, which they use to set subgoals for spatial exploration. Agents build and use a local map to predict their observations; high surprisal leads to a "fragmentation event" that truncates the local map. At these events, the recent local map is placed into long-term memory (LTM) and a different local map is initialized. If observations at a fracture point match observations in one of the stored local maps, that map is recalled (and thus reused) from LTM. The fragmentation points induce a natural online clustering of the larger space, forming a set of intrinsic potential subgoals that are stored in LTM as a topological graph. Agents choose their next subgoal from the set of near and far potential subgoals from within the current local map or LTM, respectively. Thus, local maps guide exploration locally, while LTM promotes global exploration. We demonstrate that FARMap replicates the fragmentation points observed in animal studies. We evaluate FARMap on complex procedurally-generated spatial environments and realistic simulations to demonstrate that this mapping strategy much more rapidly covers the environment (number of agent steps and wall clock time) and is more efficient in active memory usage, without loss of performance. https://jd730.github.io/projects/FARMap/

cs.AI↗

Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

We introduce Vibe-Eval: a new open benchmark and framework for evaluating multimodal chat models. Vibe-Eval consists of 269 visual understanding prompts, including 100 of hard difficulty, complete with gold-standard responses authored by experts. Vibe-Eval is open-ended and challenging with dual objectives: (i) vibe checking multimodal chat models for day-to-day tasks and (ii) rigorously testing and probing the capabilities of present frontier models. Notably, our hard set contains >50% questions that all frontier models answer incorrectly. We explore the nuances of designing, evaluating, and ranking models on ultra challenging prompts. We also discuss trade-offs between human and automatic evaluation, and show that automatic model evaluation using Reka Core roughly correlates to human judgment. We offer free API access for the purpose of lightweight evaluation and plan to conduct formal human evaluations for public models that perform well on the Vibe-Eval's automatic scores. We release the evaluation code and data, see https://github.com/reka-ai/reka-vibe-eval

cs.CL↗

Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models

We introduce Reka Core, Flash, and Edge, a series of powerful multimodal language models trained from scratch by Reka. Reka models are able to process and reason with text, images, video, and audio inputs. This technical report discusses details of training some of these models and provides comprehensive evaluation results. We show that Reka Edge and Reka Flash are not only state-of-the-art but also outperform many much larger models, delivering outsized values for their respective compute class. Meanwhile, our most capable and largest model, Reka Core, approaches the best frontier models on both automatic evaluations and blind human evaluations. On image question answering benchmarks (e.g. MMMU, VQAv2), Core performs competitively to GPT4-V. Meanwhile, on multimodal chat, Core ranks as the second most preferred model under a blind third-party human evaluation setup, outperforming other models such as Claude 3 Opus. On text benchmarks, Core not only performs competitively to other frontier models on a set of well-established benchmarks (e.g. MMLU, GSM8K) but also outperforms GPT4-0613 on human evaluation. On video question answering (Perception-Test), Core outperforms Gemini Ultra. Models are shipped in production at http://chat.reka.ai . A showcase of non cherry picked qualitative examples can also be found at http://showcase.reka.ai .

cs.CL↗

Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence

While pre-trained large-scale vision models have shown significant promise for semantic correspondence, their features often struggle to grasp the geometry and orientation of instances. This paper identifies the importance of being geometry-aware for semantic correspondence and reveals a limitation of the features of current foundation models under simple post-processing. We show that incorporating this information can markedly enhance semantic correspondence performance with simple but effective solutions in both zero-shot and supervised settings. We also construct a new challenging benchmark for semantic correspondence built from an existing animal pose estimation dataset, for both pre-training validating models. Our method achieves a PCK@0.10 score of 65.4 (zero-shot) and 85.6 (supervised) on the challenging SPair-71k dataset, outperforming the state of the art by 5.5p and 11.0p absolute gains, respectively. Our code and datasets are publicly available at: https://telling-left-from-right.github.io/.

cs.CV↗

Neuro-Inspired Fragmentation and Recall to Overcome Catastrophic Forgetting in Curiosity

Deep reinforcement learning methods exhibit impressive performance on a range of tasks but still struggle on hard exploration tasks in large environments with sparse rewards. To address this, intrinsic rewards can be generated using forward model prediction errors that decrease as the environment becomes known, and incentivize an agent to explore novel states. While prediction-based intrinsic rewards can help agents solve hard exploration tasks, they can suffer from catastrophic forgetting and actually increase at visited states. We first examine the conditions and causes of catastrophic forgetting in grid world environments. We then propose a new method FARCuriosity, inspired by how humans and animals learn. The method depends on fragmentation and recall: an agent fragments an environment based on surprisal, and uses different local curiosity modules (prediction-based intrinsic reward functions) for each fragment so that modules are not trained on the entire environment. At each fragmentation event, the agent stores the current module in long-term memory (LTM) and either initializes a new module or recalls a previously stored module based on its match with the current state. With fragmentation and recall, FARCuriosity achieves less forgetting and better overall performance in games with varied and heterogeneous environments in the Atari benchmark suite of tasks. Thus, this work highlights the problem of catastrophic forgetting in prediction-based curiosity methods and proposes a solution.

cs.AI↗

The Neumann problem on the Clifford torus in $\mathbb{S}^3$

We discuss the solution of the Neumann problem associated with the CR Yamabe operator on a subset $Ω$ of the CR manifold $\mathbb{S}^3$ bounded by the Clifford torus $Σ$. We also discuss the Yamabe-type problem of finding a contact form on $Ω$ which has zero Tanaka--Webster scalar curvature and for which $Σ$ has constant $p$-mean curvature.

math.AP↗

Redeeming Intrinsic Rewards via Constrained Optimization

State-of-the-art reinforcement learning (RL) algorithms typically use random sampling (e.g., $ε$-greedy) for exploration, but this method fails on hard exploration tasks like Montezuma's Revenge. To address the challenge of exploration, prior works incentivize exploration by rewarding the agent when it visits novel states. Such intrinsic rewards (also called exploration bonus or curiosity) often lead to excellent performance on hard exploration tasks. However, on easy exploration tasks, the agent gets distracted by intrinsic rewards and performs unnecessary exploration even when sufficient task (also called extrinsic) reward is available. Consequently, such an overly curious agent performs worse than an agent trained with only task reward. Such inconsistency in performance across tasks prevents the widespread use of intrinsic rewards with RL algorithms. We propose a principled constrained optimization procedure called Extrinsic-Intrinsic Policy Optimization (EIPO) that automatically tunes the importance of the intrinsic reward: it suppresses the intrinsic reward when exploration is unnecessary and increases it when exploration is required. The results is superior exploration that does not require manual tuning in balancing the intrinsic reward against the task reward. Consistent performance gains across sixty-one ATARI games validate our claim. The code is available at https://github.com/Improbable-AI/eipo.

cs.LG↗

Generalizing the Wythoff Array and other Fibonacci Facts to Tribonacci Numbers

In this paper, we generalize a lot of facts from John Conway and Alex Ryba's paper, \textit{The extra Fibonacci series and the Empire State Building}, where we replace the Fibonacci sequence with the Tribonacci sequence. We study the Tribonacci array, which we also call \textit{the Trithoff array} to emphasize the connection to the Wythoff array. We describe 13 new sequences.

math.NT↗

The Yamabe flow on asymptotically Euclidean manifolds with nonpositive Yamabe constant

We study the Yamabe flow on asymptotically flat manifolds with non-positive Yamabe constant $Y\leq 0$. Previous work by the second and third named authors \cite{ChenWang} showed that while the Yamabe flow always converges in a global weighted sense when $Y>0$, the flow must diverge when $Y\leq 0$. We show here in the $Y\leq 0$ case however that after suitable rescalings, the Yamabe flow starting from any asymptotically flat manifold must converge to the unique positive function which solves the Yamabe problem on a compactification of the original manifold.

math.DG↗

An Optical Parametric Amplifier via $ χ^{(2)} $ in AlGaAs Waveguides

We report parametric gain by utilizing $ χ^{(2)} $ non-linearities in a semiconductor Bragg Reflection Waveguide (BRW) waveguide chip. Under the two-mode degenerate type II phase matching, it can be shown that more than 18 dBs of parametric gain for both TE and TM modes is tenable in 100s of micrometers of device length. Polarization insensitive parametric gain can be attained within the 1550 nm region of the spectrum. These AlGaAs BRW waveguides exhibit sub-photon per pulse sensitivity. This is in sharp contrast to other types of parametric gain devices which utilize $ χ^{(3)} $, where the pump wavelength is in the vicinity of the signal wavelength. This sensitivity, which reached 0.1~photon/pulse, can usher a new era for on-chip quantum information processing using compact, micrometer-scale devices.

physics.optics↗

Using machine learning on new feature sets extracted from 3D models of broken animal bones to classify fragments according to break agent

Distinguishing agents of bone modification at paleoanthropological sites is at the root of much of the research directed at understanding early hominin exploitation of large animal resources and the effects those subsistence behaviors had on early hominin evolution. However, current methods, particularly in the area of fracture pattern analysis as a signal of marrow exploitation, have failed to overcome equifinality. Furthermore, researchers debate the replicability and validity of current and emerging methods for analyzing bone modifications. Here we present a new approach to fracture pattern analysis aimed at distinguishing bone fragments resulting from hominin bone breakage and those produced by carnivores. This new method uses 3D models of fragmentary bone to extract a much richer dataset that is more transparent and replicable than feature sets previously used in fracture pattern analysis. Supervised machine learning algorithms are properly used to classify bone fragments according to agent of breakage with average mean accuracy of 77% across tests.

cs.CV↗

Ricci Flow and Gromov Almost Flat Manifolds

We employ the Ricci flow to derive a new theorem about Gromov almost flat manifolds, which generalizes and strengthens the celebrated Gromov--Ruh Theorem. In our theorem, the condition $diam^2 |K| \leq ε_n$ in the Gromov--Ruh Theorem is replaced by the substantially weaker condition $\|Rm\|_{n/2}$ $ C_S^2 \leq \varepsilon_n$.

math.DG↗

Small curvature concentration and Ricci flow smoothing

We show that a complete Ricci flow of bounded curvature which begins from a manifold with a Ricci lower bound, local entropy bound, and small local scale-invariant integral curvature control will have global point-wise curvature control at positive times. As applications, we obtain under similar assumptions a compactness result and a gap theorem for complete noncompact manifolds with nonnegative Ricci Curvature.

math.DG↗

Sequences of the Stable Matching Problem

In this paper, we begin by discussing different types of preference profiles related to the stable marriage problem. We then introduce the concept of soulmates, which are a man and a woman who rank each other first. Inversely, we examine hell-pairs, where a man and a woman rank each other last. We generate sequences enumerating preference profiles of different types. We also calculate sequences related to the egalitarian cost, or "quality", of a matching. In total, we introduce and discuss 30 new sequences related to the stable marriage problem and discuss 6 sequences that are already in the OEIS.

math.HO↗

A $χ^{(2)}$-based AlGaAs Phase Sensitive Amplifier with Record Gain, Noise and Sensitivity

Phase sensitive amplifiers (PSAs) have the potential to empower substantial advances in emerging generations of optical communication systems as well as classical and quantum on-chip signal processing. The core building block of a PSA is a nonlinear medium. While the second-order nonlinearity ($χ^{(2)}$) is stronger than the third-order nonlinearity ($χ^{(3)}$), it is used less often in semiconductors for parametric amplification owing to the challenges of effectively phase matching the interacting waves as well as two-photon absorption of the pump. In this work, we demonstrate the successful design, fabrication, and characterization of the first $χ^{(2)}$-based semiconductor PSA using an efficient phase matching approach and a pulsed pump, based on an aluminium gallium arsenide (AlGaAs) waveguide platform. Non-centrosymmetric semiconductors such as AlGaAs offer appreciable $χ^{(2)}$. Such waveguides also achieve more than one order of magnitude greater pump field confinement when compared to other materials with large $χ^{(2)}$ such as Periodically Poled Lithium Niobate (PPLN). Our AlGaAs PSA achieves an on-chip in-phase gain, a sensitivity of 0.005 photons per pulse, and approaches theoretical minimal noise figure (NF) of 0~dB. With the capability of operating on signal states with sub-single photons per pulse, our PSA could usher in a new era of on-chip quantum circuits.

physics.optics↗

Some aspects of Ricci flow on the 4-sphere

In this paper, on 4-spheres equipped with Riemannian metrics we study some integral conformal invariants, the sign and size of which under Ricci flow characterize the standard 4-sphere. We obtain a conformal gap theorem, and for Yamabe metrics of positive scalar curvature with $L^2$ norm of the Weyl tensor of the metric suitably small, we establish the monotonic decay of the $L^p$ norm for certain $p>2$ of the reduced curvature tensor along the normalized Ricci flow, with the metric converging exponentially to the standard 4-sphere.

math.DG↗