SearcharxivSearch

arXiv subjects

Mutian Shen

Publications and source records attributed to Mutian Shen.

12 recordsLinked to original sources

AnnotateAnything: Automatic Annotation of 3D Assets for Robot Manipulation

Simulation enables scalable robot data collection, but raw 3D assets provide only geometry, lacking the semantic, interactive, and physical knowledge needed to specify where and how robots should act. In this work, we present AnnotateAnything, a general automatic annotation framework that converts passive 3D assets into manipulation-ready assets with structured, diverse, and executable manipulation labels. AnnotateAnything is built around two complementary pipelines. First, a unified visual-language annotation pipeline using vision-language reasoning to infer object semantics, interaction constraints, and 3D-grounded cues, providing human-prior guidance for identifying meaningful interaction regions. Second, a fully automatic and massively parallel physics annotation pipeline grounds these priors in each asset's geometry and physical constraints through candidate generation, geometry optimization and trajectory generation. This pipeline produces diverse and executable action annotations, including grasp poses, dexterous contacts, articulation waypoints, insertion directions, hanging affordances, and navigation targets. Using the generated annotations, we further build an asynchronous parallel simulation data-collection system across diverse objects, tasks, and robot embodiments. Experiments demonstrate that AnnotateAnything achieves superior annotation efficiency, data-collection efficiency, and task success rates over existing annotation and data-generation pipelines, while also supporting downstream tasks such as affordance detection, robotic VQA, and visual instruction finetuning. We provide project materials on the project page and plan to release the full code, annotations, and benchmark to facilitate future research. Videos, code, demo assets, and annotations are provided in supplementary materials Project page: https://tourmaline-caramel-169490.netlify.app.

cs.RO

MagicSim: A Unified Infrastructure for Executable Embodied Interaction

Robot learning and embodied agents now require simulation to serve as a shared execution substrate linking control, skills, and planning, not only as a renderer, controller testbed, or fixed task environment. Existing pipelines split these layers with "magic" actions, disconnected training environments, or forward-only renders that cannot reproduce, evaluate, and annotate the same episode. We present MagicSim, an embodied interaction infrastructure built around one deterministic batched runtime and a shared Markov decision process (MDP). From YAML-first specifications that decouple contents, placement, behavior, and agent exposure, MagicSim constructs diverse executable worlds spanning task families, interaction regimes, physics, layouts, sensors, avatars, and robot embodiments in one reset-and-step loop. A common execution interface grounds high-level commands through controllers, atomicskills, planner primitives, and asynchronous planning, realizing them as robot actions rather than simulator-side state edits. One task definition supports three capabilities: benchmark and RL evaluation, an autocollect interface that automatically turns commands into grounded trajectories, and agent/VLM-facing interaction. For automatic execution, commands flow through a Command->Skill->Planner->Robot->Record pipeline, while per-environment command, skill, planning, retry, annotation, and episode states advance independently above the shared physics tick. Successful rollouts are saved as structured multimodal trajectories aligning language supervision, action representations, visual/geometric representations, and task-level status with the executed episode. MagicSim thus unifies diverse world construction, embodied execution, task evaluation, automatic rollout generation, and interactive agent interfaces in one planner-in-the-loop runtime.

cs.RO

Pattern Expansion of Spin Glasses

We introduce a systematic method for expanding general spin-glass Hamiltonians in terms of Mattis interactions, providing a novel perspective for understanding the fundamental differences between short-range Edwards-Anderson (EA) and mean-field Sherrington-Kirkpatrick (SK) spin glasses. By iteratively extracting patterns from the coupling matrix, we expand the original spin-glass system into a Hopfield-like model (a series of Mattis interactions) plus a residual system. Our analysis reveals profound distinctions between EA and SK models: while EA models in two and three dimensions break into isolated subconnected sections after expansion, the SK model exhibits remarkable self-similar behavior, with the residual system preserving the mean-field structure and Gaussian statistics throughout the expansion process. This self-similarity manifests in exponential decay of residual matrix norms and expansion coefficients, reflecting the inherent mean-field nature of the SK model. Furthermore, we demonstrate that pattern expansion can identify ultra-low energy excitations in EA models, revealing excitations with energies that decrease rapidly with expansion step. Through connected component analysis, we quantify the size-energy relationship of these independent excitation clusters, opening new avenues for understanding the low-energy landscape of spin glasses and providing insights into the nature of metastable states.

cond-mat.dis-nn

Optimizing p-spin models through hypergraph neural networks and deep reinforcement learning

p-spin glasses, characterized by frustrated many-body interactions beyond the conventional pairwise case (p>2), are prototypical disordered systems whose ground-state search is NP-hard and computationally prohibitive for large instances. Solving this problem is not only fundamental for understanding high-order disorder, structural glasses, and topological phases, but also central to a wide spectrum of hard combinatorial optimization tasks. Despite decades of progress, there still lacks an efficient and scalable solver for generic large-scale p-spin models. Here we introduce PLANCK, a physics-inspired deep reinforcement learning framework built on hypergraph neural networks. PLANCK directly optimizes arbitrary high-order interactions, and systematically exploits gauge symmetry throughout both training and inference. Trained exclusively on small synthetic instances, PLANCK exhibits strong zero-shot generalization to systems orders of magnitude larger, and consistently outperforms state-of-the-art thermal annealing methods across all tested structural topologies and coupling distributions. Moreover, without any modification, PLANCK achieves near-optimal solutions for a broad class of NP-hard combinatorial problems, including random k-XORSAT, hypergraph max-cut, and conventional max-cut. The presented framework provides a physics-inspired algorithmic paradigm that bridges statistical mechanics and reinforcement learning. The symmetry-aware design not only advances the tractable frontiers of high-order disordered systems, but also opens a promising avenue for machine-learning-based solvers to tackle previously intractable combinatorial optimization challenges.

cond-mat.dis-nn

Neural Network Perturbation Theory (NNPT): Learning Residual Corrections from Exact Solutions

Many complex physical systems naturally decompose into an exactly solvable component augmented by a perturbative correction. Rather than directly employing neural networks to analyze complex physical systems, we introduce Neural Network Perturbation Theory (NNPT)--a correction learning approach that predicts residual perturbations after analytically subtracting known exact solutions. Using the gravitational three-body problem as testbed, we vary Jovian mass from f=0.05 to 30 times its physical value while holding network architecture fixed. An equalized-accuracy protocol with 1% tolerance reveals an unexpected non-monotonic capacity profile: capacity peaks at f=5 in the late integrable regime (3x32, 2242 parameters), remains elevated through the transition region (f~15-17), then decreases in the fully chaotic regime (f>=17, requiring only 2x32 with 1186 parameters)--a 47% reduction from peak. With symplectic integrator energy conservation below 2x10^{-4}, this counterintuitive phenomenon reflects genuine physical structure rather than numerical artifacts. Sequential correction experiments show negligible refinement (||y2||/||y1||~0.997), confirming single-stage networks capture dominant perturbative features without hierarchical decomposition. The capacity transition at f_c=16.6+-2.8 aligns with Chirikov's resonance-overlap criterion. Intermediate-complexity regimes impose maximal capacity requirements, while fully chaotic dynamics undergo ergodic smoothing--trajectory-specific fluctuations become irreducible noise, leaving only statistically smooth corrections requiring fewer parameters.

physics.comp-ph

The Physics of Local Optimization in Complex Disordered Systems

Limited resources motivate decomposing large-scale problems into smaller,``local" subsystems and stitching together the so-found solutions. We explore the physics underlying this approach and discuss the concept of ``local hardness", i.e., the complexity of predicting local properties of the solution from local information, for the ground-state problem of both P- and NP-hard spin-glasses and related frustrated spin systems. Depending on the model considered, we observe varying scaling behaviors in how errors associated with local predictions decay as a function of the size of the solved subsystem. These errors are intimately connected to global critical threshold instabilities, characterized by gapless, avalanche-like excitations that follow scale-invariant size distributions. Away from criticality, local solvers quickly achieve high accuracy, aligning closely with the results of the computationally much more expensive global minimization. We leverage these findings to introduce a heuristic contraction-based algorithm for globally studying spin-glass ground states. The local solvers further display sharp imprints of the phase transition from the spin-glass to the ferromagnetic phase as the distribution of spin-glass couplings is shifted, as well as characteristic differences for the infinite-range model, implying the existence of specific classes of local hardness. Our findings shed light on how Nature may operate solely through local actions at her disposal.

cond-mat.dis-nn

The Eggbox Ising Model

We introduce the Eggbox Ising model, a tunable construction of rugged energy landscapes defined by distances to a prescribed set of patterns. Correlated pattern ensembles realize arbitrary k-step replica-symmetry-breaking structures and controllable Parisi overlap distributions p(q), consistent with the hierarchical overlap structure observed in a simple word-embedding example from empirical data. A softened variant allows a systematic expansion leading to Hopfield-type couplings (and higher-body terms). We analyze the density of states and show that suitable potentials induce discontinuous finite-temperature transitions with metastability and hysteresis.

cond-mat.stat-mech

Transform then Explore: a Simple and Effective Technique for Exploratory Combinatorial Optimization with Reinforcement Learning

Many complex problems encountered in both production and daily life can be conceptualized as combinatorial optimization problems (COPs) over graphs. Recent years, reinforcement learning (RL) based models have emerged as a promising direction, which treat the COPs solving as a heuristic learning problem. However, current finite-horizon-MDP based RL models have inherent limitations. They are not allowed to explore adquately for improving solutions at test time, which may be necessary given the complexity of NP-hard optimization tasks. Some recent attempts solve this issue by focusing on reward design and state feature engineering, which are tedious and ad-hoc. In this work, we instead propose a much simpler but more effective technique, named gauge transformation (GT). The technique is originated from physics, but is very effective in enabling RL agents to explore to continuously improve the solutions during test. Morever, GT is very simple, which can be implemented with less than 10 lines of Python codes, and can be applied to a vast majority of RL models. Experimentally, we show that traditional RL models with GT technique produce the state-of-the-art performances on the MaxCut problem. Furthermore, since GT is independent of any RL models, it can be seamlessly integrated into various RL frameworks, paving the way of these models for more effective explorations in the solving of general COPs.

cs.LG

Universal fragility of spin-glass ground-states under single bond changes

We consider the effect of perturbing a single bond on ground-states of nearest-neighbor Ising spin-glasses, with a Gaussian distribution of the coupling constants, across various two and three-dimensional lattices and regular random graphs. Our results reveal that the ground-states are strikingly susceptible to such changes. Altering the strength of only a single bond beyond a critical threshold value leads to a new ground-state that differs from the original one by a droplet of flipped spins whose boundary and volume diverge with the system size -- an effect that is reminiscent of the more familiar phenomenon of disorder chaos. These elementary fractal-boundary zero-energy droplets and their composites feature robust characteristics and provide the lowest-energy macroscopic spin-glass excitations. Remarkably, within numerical accuracy, the size of such droplets conforms to a nearly universal power-law distribution with exponents dependent on the spatial dimension of the system. Furthermore, the critical coupling strengths adhere to a stretched Gaussian distribution that is predominantly determined by the local coordination number.

cond-mat.dis-nn

Reply to: Deep reinforced learning heuristic tested on spin-glass ground states: The larger picture

We wish to thank Stefan Boettcher for prompting us to further check and highlight the accuracy and scaling of our results. Here we provide a comprehensive response to the Comment written by him. We argue that the Comment did not account for the fairness of the comparison between different methods in searching for the spin-glass ground states. We demonstrate that, with a reasonably larger number of initial spin configurations, our results agree with the asymptotic scaling form assumed by finite-size corrections.

cond-mat.dis-nn

Finding spin glass ground states through deep reinforcement learning

Spin glasses are disordered magnets with random interactions that are, generally, in conflict with each other. Finding the ground states of spin glasses is not only essential for the understanding of the nature of disordered magnetic and other physical systems, but also useful to solve a broad array of hard combinatorial optimization problems across multiple disciplines. Despite decades-long efforts, an algorithm with both high accuracy and high efficiency is still lacking. Here we introduce DIRAC - a deep reinforcement learning framework, which can be trained purely on small-scale spin glass instances and then applied to arbitrarily large ones. DIRAC displays better scalability than other methods and can be leveraged to enhance any thermal annealing method. Extensive calculations on 2D, 3D and 4D Edwards-Anderson spin glass instances demonstrate the superior performance of DIRAC over existing methods. As many hard combinatorial optimization problems have Ising spin glass formulations, our results suggest a promising tool in solving these hard problems. Moreover, the presented algorithm will help us better understand the nature of the low-temperature spin-glass phase, which is a fundamental challenge in statistical physics.

cond-mat.dis-nn

Compressed Sensing by Shortest-Solution Guided Decimation

Compressed sensing is an important problem in many fields of science and engineering. It reconstructs signals by finding sparse solutions to underdetermined linear equations. In this work we propose a deterministic and non-parametric algorithm SSD (Shortest-Solution guided Decimation) to construct support of the sparse solution under the guidance of the dense least-squares solution of the recursively decimated linear equation. The most significant feature of SSD is its insensitivity to correlations in the sampling matrix. Using extensive numerical experiments we show that SSD greatly outperforms L1-norm based methods, Orthogonal Least Squares, Orthogonal Matching Pursuit, and Approximate Message Passing when the sampling matrix contains strong correlations. This nice property of correlation tolerance makes SSD a versatile and robust tool for different types of real-world signal acquisition tasks.

eess.SP