Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Are Coreset Selection Methods Worth Their Cost?

Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmark that standardizes downstream training and charges selection and training to the same auditable wall-clock budget, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, with over 1,500 released runs. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies. Our two budget studies test whether that verdict survives when every selector is granted its most favorable operating point. Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In fixed-budget duels on ImageNet-1K, training on all data for fewer epochs beats every selection strategy we probe while also costing the least. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further quantify when selection does pay back through subset reuse, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by nearly 6 points. Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.

cs.LG↗

The Requirement of (at least) Complex Structure for Quantum Mechanics

It is argued that many real-valued constructions of quantum mechanics are only \emph{nominally} real in that the operators and states are restricted to impose a \emph{complex structure} on the Hilbert space. That is, complex algebra between pairs of elements representing complex numbers is preserved in these formulations. It is therefore mistaken to think of these constructions as being `real' as they are actually a representation of complex linear algebra. It is then shown that under the assumptions of state normalisation and strict energy conservation, this complex structure (at the very least) is required to allow for time evolution in quantum mechanics. This is only a \emph{minimum} requirement, as our arguments do not preclude the formulation of quantum mechanics in terms of hyper-complex entities, such as quaternions.

quant-ph↗

Improved upper bound on the number of distinct k-decks for any k and alphabet size by counting the independent parameters

Data stored in synthetic DNA is retrieved by shotgun sequencing, which returns short subsequences rather than the stored word itself. A natural abstraction of this readout is the $k$-deck of a word: the vector recording how often each word of length $k$ occurs as a subsequence. Two stored words are distinguishable from their readouts exactly when their $k$-decks differ, so the number $D_{q,k}(n)$ of distinct $k$-decks of words of length $n$ over an alphabet of size $q$ measures what a length-$k$ readout retains. We analyse the degrees of freedom remaining in a $k$-deck once all shorter decks are fixed. Within each class of words having prescribed letter multiplicities, the length-$k$ entries are confined to an affine subspace whose dimension is exactly the number of Lyndon words with the same multiplicities, which we give in closed form as a Möbius sum. Writing $L_q(j)$ for the number of Lyndon words of length $j$ over an alphabet of size $q$, we deduce the improved upper bound \[ D_{q,k}(n)=O\!\left(n^{E_q(k)}\right),\qquad E_q(k)=\sum_{j=1}^{k}j\,L_q(j)-1 . \] In the case of a binary alphabet this bound satisfies $D_{2,k}(n)=O\!\left(n^{4\cdot 2^{k-1}}\right)$. We then prove matching lower bounds in the first two nontrivial cases: $D_{q,2}(n)=Θ\!\left(n^{q^2-1}\right)$ for every alphabet size $q$, and $D_{2,3}(n)=Θ(n^{9})$ for the binary alphabet. The latter confirms, for $q=2$ and $k=3$, our conjecture that the upper bound has the correct degree for every $q$ and $k$.

math.CO↗

Perfect Codes in the Johnson Scheme Hardly Exist

In his pioneer work from 1973, Delsarte conjectured that there are no nontrivial perfect codes in the Johnson scheme J$(n,w)$. While in most other important schemes the existence problem for perfect codes was settled, the problem is still open in the Johnson scheme. In this work we considerably reduce the possible existence of such codes. We prove that there are no $e$-perfect codes in the Johnson scheme when $e \not\in \{1,2,4,9,10,12,16\}$. These seven cases will be considered and solved in a follow up paper.

math.CO↗

Bayesian Deck-of-cards-based Ordinal Regression with Sequential Preference Elicitation

The Deck-of-cards-based Ordinal Regression (DOR) infers a value function from a ranking of reference alternatives in which the Decision Maker (DM) inserts blank cards between consecutive levels to express preference intensity. DOR, and its stochastic extension (SMAA-DOR), treat these answers as hard constraints defining a set of compatible value functions. We propose B-DOR, a probabilistic reformulation of DOR in which each pair of adjacent levels yields an ordinal observation, the declared direction and the number of cards, modelled through a cumulative-link likelihood that relates the number of blank cards to the latent value difference between alternatives. Two Bayesian inference algorithms are proposed: BAYES-DOR samples the whole posterior distribution by Hamiltonian Monte Carlo; FTRL-DOR tracks the maximum a posteriori estimate by constrained convex optimization. Moreover, through a multi-step elicitation process, elicitation can be spread over several short sessions reducing the cognitive burden on the DM. Both algorithms enjoy logarithmic regret bounds for prediction that hold for any sequence of DM responses and that guide the choice of the prior hyperparameters. A Monte Carlo study over 768 configurations shows that accuracy grows with the number of sessions, that blank cards add significant information over preference directions alone, that both algorithms maintain good performance under inconsistent answers, and that both outperform DOR and SMAA-DOR. An illustrative application to Italian regional healthcare performance demonstrates the practical applicability of the approach for building composite indicators.

stat.ML↗

HOIBlender: Blending Lightweight Detection with Vision-Language Priors for Efficient Human-Object Interaction Detection

Human-object interaction (HOI) detection requires grounding an interacting human-object pair and recognizing the verb that links them, often under severe long-tail supervision. Recent methods improve accuracy with stronger detectors and vision-language priors, but many still stack heavy transformer encoders, intricate denoising schedules, or post-hoc semantic calibration on top of the detector. We present \textbf{HOIBlender}, an efficient HOI detector named after its core design principle: blending detector-grounded visual tokens, spatial subject-object reasoning, and BLIP-2 semantic priors inside one lightweight decoding pipeline. HOIBlender builds on an RF-DETR/LW-DETR-style foundation with a DINOv2 backbone and selects top-$K$ image-conditioned tokens directly from the multi-scale projector as subject and object candidates, removing the dedicated encoder stage retained by prior HOI methods. A dual-stage decoder first stabilizes human-object geometry and then performs verb and HOI classification through progressive BLIP-2 prior fusion, with classifier weights initialized from BLIP-2 text embeddings for long-tail categories. Grouped-query training further enriches optimization without increasing inference cost. Across three model scales (Nano, Small, 2XL), HOIBlender consistently outperforms SOV-STG-VLA and Hybrid-SOV-VLA on HICO-DET, reaching $44.49$ Default Full mAP in only $9$ training epochs while maintaining competitive latency and parameter budgets. These results show that lightweight detection, structured spatial-semantic decoding, and deeply integrated vision-language priors can be blended into a single efficient HOI pipeline.

cs.CV↗

Receding-Horizon Pushing with Composable Object-Centric Policies

Non-prehensile manipulation is practical for relocating large, heavy, or geometrically ungraspable objects. Yet, long-horizon pushing of arbitrarily-shaped 3D objects couples three problems: 1) where to push the object so as to approach the target pose, 2) whether each push is stable and reachable, 3) whether subsequent actions remain feasible. We present an object-centric pushing policy within a feedback-guided hierarchical framework. At the low level, a learning-based policy predicts contact actions from a pose- and scale-normalized point cloud, conditioned on a near single-step subgoal. A stability score is applied to evaluate the predicted contacts by a quasi-static sliding-versus-tipping analysis. At the high level, BIT$^*$ first searches for an object path, and the next several subgoals are checked by contact prediction and robot motion planning for future feasibility. Failed motion plans, as feedback, change the local path costs and trigger re-planning. During execution, only the first feasible action is executed. In simulation, we evaluate 22 objects in six different scenes, upon which we also conduct comprehensive ablation studies. Results demonstrate that our method outperforms baselines with a clear margin and can reliably achieve long-horizon object pushing tasks under different situations. We also report quantitative real-robot experiments with a Franka arm and qualitative demonstrations with a mobile manipulator for large and heavy objects, with directly zero-shot sim-to-real transfer.

cs.RO↗

Homological critical points for hypergraphs and simplicial complexes

In this paper, we study the homological critical points for persistence hypergraphs and persistence simplicial complexes. With the help of the relative homology groups, we define the homological critical points for persistence morphisms between persistence hypergraphs and persistence simplicial maps between persistence simplicial complexes. We prove some commutative diagrams of subset relations for the homological critical points of persistence hypergraphs as well as persistence morphisms between them. We also prove some commutative diagrams of subset relations for the homological critical points of persistence simplicial complexes as well as persistence simplicial maps between them. As examples, we use the parametric configuration spaces to construct the parametric independence complexes as the space of parametric packings and construct the parametric dominating hypergraphs as the space of parametric coverings.

math.AT↗

TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation

Reactive vision-language-action (VLA) policies suffer from task-state aliasing in long-horizon manipulation, where identical multimodal inputs call for distinct, context-dependent actions. Given that pretrained VLAs already possess rich control primitives to express diverse behaviors, we hypothesize that the execution bottleneck lies not in policy capacity, but in input ambiguity. In this paper, we propose TaskAnchor, a lightweight adapter that grounds task state by injecting execution context into the VLA's native input space. During post-training, TaskAnchor learns to represent the semantic execution stage as a milestone-supervised coordinate prepended to the language instruction, while incorporating fine-grained historical evidence via a residual update to the current visual tokens. This formulation avoids generating complex subtask instructions and leaves the backbone architecture unchanged. Across long-horizon benchmarks, TaskAnchor delivers substantial gains, achieving approximately 6 times the average success rate of the pi0.5 and X-VLA baselines on RMBench and more than doubling the task success rate of pi0.5 on RoboMemArena. Real-robot experiments further validate reliable multi-stage execution, with the same policy adapting its subsequent behaviors using earlier human interactions as in-context cues. Our project website is available at https://taskanchor.netlify.app/.

cs.RO↗

Tail-Weight Control and Localized Generalization in Nearly Low-Rank Adversarial Classification

Empirical ramp fitting can assign weight to pure-noise features even when the population optimum ignores them. We quantify this gap for norm-constrained adversarial classification with Gaussian signal and noise. The variance cost relative to normalized signed mean separates into two factors: selecting observations inside the active margin window and the curvature induced by the norm constraint. Changing the tail variance leaves the activewindow probability unchanged but changes the second factor. With positive attack budget and a signal-only predictor of risk below one half, we prove a uniform quadratic tail-deletion bound, including at zero tail variance. Sufficiently accurate approximate global empirical minimizers admit exact fixeddimensional asymptotic covariances in the low-risk regime with isotropic principal covariance. For positive tail variance at most principal variance, the product exceeds one; an additional moment condition transfers it to expected excess ramp and robust classification risks. A wide window analysis characterizes when this ordering reverses. Controlled experiments test the decomposition, and a separate contamination study examines its scope outside the Gaussian training model.

cs.LG↗

Matching Rules for a Three-Dimensional Strongly Aperiodic Monotile

A recent pre-print [arXiv:2609.19214] proposed a three-dimensional (3D) strongly aperiodic monotile: a shape that tiles Euclidean space only non-periodically and which admits no symmetry of infinite order. The proof takes the 3D Chair tile identified previously by Lee and Moody, and adds geometric decorations to the faces so as to force non-periodicity (without these decorations The Chair also admits periodic tilings). Here we establish general requirements on face decorations to achieve the same end, in order to facilitate the search for physical realisations. We find that the requirements are minimal. We provide matching rules using three colours of arrow that are equivalent to the original rules, in that they force the same local and global configurations. These rules force Chairs to compose into `Superchairs' with doubled linear dimensions. In this process the matching rules themselves compose uniquely. We find that the same global structure can be forced using simpler rules based on the colours of squares, regardless of orientation. Any physical system encoding these rules (geometrically or otherwise) will force the same strongly aperiodic monotilings. We provide simple examples.

cond-mat.other↗

A Physics-Conditioned Neural Operator for Generalization of Atrioventricular Valve Mechanics across Pressure and Tissue Properties

Mitral and tricuspid regurgitation are the most common regurgitant valvular lesions, yet only a minority of severe cases undergo corrective surgery. Rapid assessment of valve mechanics could enable earlier, more precise intervention, but finite element (FE) analysis is slow to repeat across the many loading and tissue-property values of interest, which for a given valve are not known in advance. We introduce the Physics-Conditioned Neural Operator (PCNO), a transformer-based neural operator predicting leaflet displacement, strain, and stress fields conditioned on systolic blood pressure and tissue properties. PCNO is trained on, and evaluated against, FEBio simulations of functional and regurgitant mitral and tricuspid valves and of three mitral pathologies. With pressure and all material parameters simultaneously outside the training support, displacement error against FE reaches 4.48% and the mean errors of unsupervised geometric measures of valve function stay within 3.5%, indicating a conditioned solution operator over parameter space rather than an interpolator of the training set. Under this shift, PCNO is also more accurate than graph neural network and graph neural operator baselines trained on the same simulations, with the largest margins in stress.

cs.LG↗

Bulk-edge sticking beyond the Perron mode in Gaussian softmax attention

We study row-softmax self-attention with independent Gaussian query and key weights in the proportional regime, at fixed inverse temperature. Hayase, Collins, and Karakida proved Gaussian equivalence for the empirical squared singular-value distribution after removal of the Perron direction. A global law alone does not exclude finitely many nonleading outliers. We prove that no such outliers persist: the rescaled squared singular value $\ell s_k(A)^2$ converges in probability to the upper edge of their bulk law for every fixed $k\ge 2$. In fact, this convergence is uniform over any deterministic sublinear number of leading non-Perron indices. The proof uses an exact decomposition of the softmax normalization, conditions on the key matrix, identifies the conditional covariance exactly with a diagonally conjugated inner-product kernel, linearizes that kernel in operator norm, and applies the outside-support local law of Fan, Ma, Paquette, and Wang. A separate stability argument identifies the finite conditional deformed Marchenko-Pastur edge with the limiting bulk edge. We also derive a scalar formula for that edge throughout the proportional regime and recover the explicit square-model formula of Hayase, Collins, and Karakida, including its physical branch.

math.PR↗

Binding-Motivated Contextuality: A Cross-Domain Cyclic Test in Perception and Judgment

Psychophysics and decision research study perceptual binding and judgment contextuality apart. We argue both are scored against the same cyclic noncontextuality inequalities and share one convex global-consistency geometry -- not one cohomology class -- though only contextuality is tested, since binding's own residual vanishes here. Building on sheaf formulations of predictive coding (Seely 2025) and contextuality (Abramsky & Brandenburger 2011), a cyclic set of pairwise judgments admits a noncontextual explanation exactly when the cyclic (Suppes-Zanotti or n-cycle) inequalities hold, and its severity is measured by the complete contextual fraction CF (Abramsky, Barbosa & Mansfield 2017) rather than by the Cech invariant, which can miss it (Caru 2017). We build the perceptual arena from two binary judgments per cyclic-dominance pairing, scored against the same inequalities as the survey; an appendix rejects the one-bit alternative on construct-validity grounds. The central test is cross-domain: one cohort performs both arenas, and a shared latent tolerance predicts a positive correlation between their signed cyclic margins V*, the pre-clamp quantities behind CF. Both are inconsistency scores, so general response consistency confounds a bare correlation, and the prediction is therefore confound-residualized against a variance- and reliability-matched control. A positive result would support a shared residual association, not a common causal mechanism, which we prove this design cannot identify at any sample size. All three tests are designed but unrun, and the perceptual arena is a proposed instantiation with pilot gates -- a methodological proposal, not a confirmatory report.

q-bio.NC↗

Object-Centric Conditioning for Visuomotor Flow Matching

Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical motion patterns and entangled global visual representations can jointly reduce robustness under spatial out-of-distribution (OOD) shifts and visual distractors. In this work, we propose SlotFlow, an object-centric flow matching policy for robust visuomotor manipulation. SlotFlow decouples scene observations into semantic ("what") features and lightweight image-plane spatial ("where") cues to provide object-aware policy conditioning and current-state grounding. The semantic representation suppresses irrelevant background correlations, while the spatial cue improves adaptation to shifted object configurations. Extensive simulation and real-world experiments demonstrate improved robustness under visual distractors and severe spatial perturbations while preserving the low-step inference efficiency of A2A. Controlled initialization and perception ablations further identify object-centric grounding as a major source of the gains and show that it complements, rather than replaces, useful historical motion priors.

cs.RO↗

ME-Brain-1.0: Memory, Cognition and Action for Evolving Embodied Intelligence

Current embodied systems largely rely on pretrained capabilities that remain fixed after deployment, limiting their ability to learn from physical interaction. We introduce MachEmbodied-Brain (ME-Brain), a self-evolving embodied system organized around a closed loop of action execution, experience acquisition, experience evolution, and improved execution. Evolvable Memory consolidates multimodal trajectories into hierarchical, reusable experience; Cognitive Core transforms physical experience into transferable skills; and the Action Model combines event-driven keyframes, EventCell local-world prediction, and action-conditioned memory modulation to focus computation on decision-critical moments, regions, and historical evidence. Together, these modules shift embodied intelligence from train-and-freeze to deploy-and-evolve without model retraining. Cognitive Core outperforms the strongest comparison models by 8.2 and 9.6 points on embodied and agent benchmarks. The Action Model achieves 47.88% mean success on RoboMME, a 3.26-point improvement over the strongest baseline. On RoboDojo, it reaches a 21.51 mean Score and 16.03% success rate, exceeding $π_{0.5}$ by 10.10 and 9.12 points. On the six-task ME-RealBench, ME-Brain achieves a 69.5 mean Score and 66.7% success rate, outperforming DM0.5 by 12.8 and 11.7 points, respectively.

cs.RO↗

Statistical mechanics of multipartite entanglement in hypergraph states

We investigate multipartite entanglement in a particular family of pure $n$-qubit hypergraph states through a statistical-mechanics framework, where the average bipartite purity maps onto an effective Hamiltonian of $2^n$ classical binary spins. In this correspondence, each hypergraph state uniquely corresponds to a classical spin configuration, while temperature serves as a control parameter that continuously interpolates between a uniform ensemble of random hypergraph states at high temperature and maximally multipartite entangled states (MMES) at zero temperature. Remarkably, the exponential of the zero-temperature entropy directly gives the number of MMES within the set of hypergraph states. For small system sizes ($n \leq 5$), we perform an exact enumeration, fully characterizing the energy landscape and associated thermodynamic observables, and validating known MMES counts. For larger systems ($n = 6$ and $7$), where exact methods become computationally infeasible, we employ simulated annealing and parallel tempering algorithms to efficiently sample the exponentially large state space. Our analysis yields quantitative predictions of the number of MMES and reveals how entanglement is statistically distributed across the sets of hypergraph states. These results establish hypergraph states as an ideal platform for investigating multipartite entanglement through thermodynamic methods, offering both computational advances and physical insights into the structure of quantum entanglement in restricted families of quantum states.

quant-ph↗