SearcharxivSearch

arXiv subjects

Pengfei Li

Publications and source records attributed to Pengfei Li.

At least 19 recordsLinked to original sources

CAP: Continuously Adaptive Perception-Blind Humanoid Locomotion via Learned Denoising

Humanoid locomotion across complex terrain demands forward-looking exteroception to anticipate obstacles, yet this signal is unreliable in real-world deployment, failing partially and intermittently. Existing perceptive policies often assume that depth observations remain clean and in-distribution, while recent attempts to unify perceptive and blind control typically route or switch between separate sub-policies, leaving recoverable information in partially corrupted depth unexploited. We instead propose CAP, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information. A coupled training recipe pairs a depth-noise curriculum on the world-model input with world-model feature dropout on the policy-facing latent, exposing the policy to failures across the entire perception-quality spectrum. In simulation, CAP matches or improves upon perceptive baselines when depth remains informative, and degrades more smoothly than a binary-switching baseline as perception worsens. On the Unitree G1, controlled trials and indoor-outdoor deployments demonstrate perception-robust locomotion under intermittent occlusion, real-sensor corruption, and outdoor depth artifacts.

cs.RO

Semiparametric Receiver Operating Characteristic Analysis in the Presence of an Imperfect Reference Standard via a Box-Cox Density Ratio Model

Receiver operating characteristic (ROC) analysis is commonly used to evaluate the diagnostic accuracy of continuous biomarkers. In practice, the true disease status may be unavailable and only a nominal disease status provided by an imperfect reference standard is observed. Existing nonparametric methods have been developed for ROC analysis in this setting, but may suffer from reduced estimation efficiency, numerical instability, or sensitivity to the choice of biomarker scale. We propose a semiparametric method based on a Box-Cox density ratio model, which links the biomarker distributions of the truly healthy and diseased populations while leaving the baseline distribution unspecified. A key feature of the proposed method is that the transformation parameter is estimated from the data rather than specified in advance, allowing the density-ratio structure to adapt to different transformation scales. We develop an empirical likelihood approach for estimation and an expectation-maximization algorithm for computation. We establish the asymptotic distributions of estimators of the ROC curve, area under the curve, Youden's index, and the sensitivity and specificity at the Youden-optimal cutoff, and develop bootstrap confidence intervals and a goodness-of-fit test. Simulation studies demonstrate that the proposed method provides accurate and numerically stable estimation and inference across a range of distributional settings without requiring the transformation scale to be specified in advance. The proposed method is illustrated using data from a malaria study.

stat.ME

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

Multimodal Large Language Models (MLLMs) show strong progress on vision-language tasks, yet their reliability in safety-critical settings remains underexplored. Fire-smoke understanding is central to public safety and disaster response, but most existing benchmarks lack diverse real-world scenarios and context-aware evaluation. We introduce SAFIRE, a large-scale benchmark for fire-smoke understanding in MLLMs, comprising 83K captioned images from 20 scenarios and 193K multiple-choice VQA (MCVQA) generated from a 9.7K-image subset, spanning 10 evaluation dimensions from basic perception to higher-order reasoning. A GPT-5.4-assisted multi-stage verification pipeline with MLLM majority voting ensures annotation quality. Evaluating ten open-source MLLMs (8B-38B) yields an average accuracy of 61.9%, exposing major gaps in safety-critical reasoning. We further show that adapting vision encoders with only 7% of our domain-specific data boosts fire-scene classification accuracy from 20.1% to 64.5%, indicating that carefully curated data can yield substantial gains even when data volume is limited. All datasets, models, and code are available at https://risys-lab.github.io/SAFIRE/.

cs.CV

Revisiting dependence in multiple testing: empirical distribution approaches for FDP control

Large-scale multiple hypothesis testing is central to the analysis of high-throughput data, where controlling false discoveries is critical. Classical procedures typically rely on theoretical null distributions and often adjust for dependence among test statistics, but these approaches may be misleading when the empirical distribution of null statistics deviates from theoretical assumptions. Motivated by this observation, we investigate the role of the empirical cumulative distribution function (c.d.f.) of null test statistics in controlling the false discovery proportion (FDP). We first show that, under an oracle scenario where the empirical c.d.f. of the test statistics for all null hypotheses is known, FDP control can be achieved optimally regardless of the dependence structure, highlighting that explicit modeling of dependence may be unnecessary. Building on this insight, we propose an empirical c.d.f.-based FDP control (eFDP) method, implemented via a multivariate mixture model framework and a nonparametric estimation procedure for the empirical c.d.f.s, establish its asymptotic convergence, and construct an FDP control procedure that achieves asymptotic FDP control. Extensive simulations demonstrate that eFDP attains more accurate FDP control and higher power than existing approaches, particularly under strong dependence, and analysis of a high-dimensional breast cancer gene expression dataset confirms its practical utility.

stat.ME

Influence of twist direction and large deformation on soft material torsional contact

Shear-induced contact area reduction is widely observed in soft contacts, yet recent torsional experiments have revealed a more complex non-monotonic evolution in which the contact area first increases and then decreases with twist angle. The mechanism responsible for this initial area increase and the role of large deformation in the overall area evolution remain unclear. In this study, we experimentally investigate the torsional contact response of soft Polydimethylsiloxane (PDMS) spheres by combining forward-backward twist tests with a systematic variation of the curing-agent-to-base ratio to tune material softness and deformation level. The loading-unloading tests show that the torsional interface is strongly irreversible: during unloading, the contact area follows a decrease-increase-decrease path rather than retracing the loading branch, and repeatable petal-like edges appear, indicating a wrinkle-induced surface instability. By decreasing the mixing ratio, we find that larger deformation strengthens the area-reduction contribution and eventually suppresses the initial area increase, leading to a monotonic area decrease during loading for sufficiently soft PDMS. Softer PDMS also exhibits lower shear strength, weaker torque oscillations, and improved repeatability. The results provide experimental evidence that large deformation can drive shear-induced contact area reduction, while the origin of the initial area increase remains unresolved. These findings narrow the possible mechanisms (e.g., triboelectrification) responsible for the initial area increase and provide a stringent benchmark for frictional contact models of soft interfaces.

cond-mat.soft

Quiescent and traveling solitons in the fractional parametrically driven damped nonlinear Schr\"{o}dinger equation

We systematically investigate the existence, stability, and dynamics of optical solitons in the framework of the one-dimensional nonlinear Schr\"{o}dinger equation with the Riesz-fractional diffraction operator, cubic self-focusing, and linear loss, balanced by a linear parametric drive. The model, which can be realized in a laser cavity, produces standing and moving solitons, the latter ones existing below a critical velocity. One of the soliton species is stable in a wide range of parameters, while others are unstable. The fractional diffraction significantly alters the existence conditions and stability thresholds of the solitons. Collision between moving solitons are considered too. The results essentially expand the variety of nonlinear modes in media with fractional diffraction.

nlin.PS

FedSPM: Routing-Enabled Federated Learning under Dual Heterogeneity via Semiparametric Mixture

Routing-prediction federated learning has emerged as a new paradigm that reframes inter-client heterogeneity as a resource for system-level intelligence: at inference time, the server routes each external query to the best-matched client for prediction. Existing approaches, however, typically treat each client as internally homogeneous, overlooking latent subpopulations within local data. For example, patients with the same diagnosis at one hospital may exhibit morphologically distinct disease subtypes. The coexistence of inter-client and intra-client heterogeneity, which we call dual heterogeneity, can impair both routing and prediction. To address this challenge, we propose FedSPM, a routing-enabled semiparametric mixture framework that represents each client using client-specific latent components. Each component combines a predictive distribution for classification with a feature distribution for routing. To flexibly model feature distributions while effectively sharing information across clients, FedSPM models their density ratios relative to a common nonparametric measure estimated via empirical likelihood. We develop a federated expectation-maximization algorithm that optimizes a tractable surrogate and prove convergence of the exact profiled objective at the standard $\mathcal{O}(1/\sqrt{T})$ rate when the surrogate errors are properly controlled. Experiments on controlled benchmarks and real-world medical data demonstrate consistent improvements in routing and prediction under dual heterogeneity. Code is available at https://github.com/zijianwang0510/FedSPM.

cs.LG

Legible Shared Autonomy: Implicit Communication of Robot Belief through Motion

Shared autonomy systems combine user input with autonomous assistance to help users with motor impairments control robot arms to perform everyday manipulation tasks, by inferring user goals and providing appropriate guidance. However, the robot's internal beliefs about user goals cannot be observed by users. Traditional shared autonomy systems provide assistance along efficient shortest paths toward inferred goals, but when multiple objects lie in similar directions, such assistive motion remains ambiguous and fails to reveal the specific goal identified by the robot. This creates two critical problems. First, when the robot correctly infers the goal, users continue controlling because they cannot perceive understanding from ambiguous assistive motion, wasting effort when autonomous completion would suffice. Second, when the robot misunderstands intent, users cannot quickly detect errors until assistive motion diverges significantly, requiring substantial corrective input. We address this by introducing legible motion into shared autonomy, where robot actions must both advance toward the goal and clearly reveal which goal has been inferred, enabling users to understand the robot's beliefs and adjust control accordingly. The robot modulates communication strength through confidence-aware adaptive authority allocation by providing assertive legible assistive actions when confident while increasing user authority when uncertain, transforming shared autonomy into transparent bidirectional collaboration. User studies including simulation, and physical experiments with a six-degree-of-freedom robot arm demonstrate that legible shared autonomy significantly improves users' understanding of robot beliefs and reduces user control effort compared to standard shared autonomy.

cs.RO

Analytical and fitting formulae for solutions to Lyman-alpha radiative transfer equations: the effects of geometry, recoil, and velocity gradients

Lyman-alpha (Ly$\alpha$) radiative transfer (RT) is important in many astrophysical environments and governed by multiple physical processes. In this paper, we provide analytical formulae/procedures for the solutions to Ly$\alpha$ RT equations under three simple geometrical symmetries and investigate the effects of atomic recoil and gas bulk motion. We first study Ly$\alpha$ spectra by solving Ly$\alpha$ RT equations for a static, uniform gas cloud under cylindrical geometry. The solution is verified through Ly$\alpha$ Monte Carlo RT simulations, and compared to those under slab and spherical geometries in literature. Second, to characterise the recoil effect, we empirically modify recoil-free Ly$\alpha$ spectra. The method is motivated by Ly$\alpha$ RT equations with recoil and justified by simulations. Finally, we account for constant velocity gradients in Ly$\alpha$ RT equations and obtain series solutions for Ly$\alpha$ spectra. The solutions demonstrate good agreement to Ly$\alpha$ spectra from simulations for small velocity gradients (i.e. edge velocity $v_{\rm E}$ of a cloud being comparable to the thermal velocity $b$) but become less accurate for large ones. To characterise Ly$\alpha$ spectra under large velocity gradients, we empirically extend the functional form of solutions and constrain them from fitting simulated Ly$\alpha$ spectra. The resulting fitting formulae show significant improvement for large velocity gradients ($v_{\rm E}/b \sim 100$) under large optical depths. The analytical study of Ly$\alpha$ spectra in this work completes the set of solutions under simple geometries, provides physical insights for Ly$\alpha$ RT under recoil and velocity gradient, and develops analytical tools for theoretical studies that require inputs from Ly$\alpha$ RT.

astro-ph.GA

Fed-CausalDiff: Decoupled Synchronization for Federated Do-Simulation and Policy Evaluation

While federated learning enables collaborative modelling on decentralised data, standard methods merely fit historical observations. This purely observational approach is fundamentally insufficient for interventional inference and policy evaluation, as sequential actions dynamically alter future states. We propose \textbf{Fed-CausalDiff}, a federated causal diffusion framework for do-simulation. The architecture decomposes the evolution of the latent state into a global causal score function and a local confounding score function. This design enables \emph{decoupled synchronisation} (DSS), where clients aggregate only the shared causal mechanism while retaining site-specific confounders locally to handle heterogeneity. Experiments on four datasets demonstrate that Fed-CausalDiff achieves better ATE and policy-value estimation accuracy, offering a favorable trade-off between communication cost and inference fidelity.

cs.LG

On Revisiting Entropy for Identifying Mislabeled Images

Mislabeled samples in training datasets severely degrade the performance of deep networks, as overparameterized models tend to memorize erroneous labels. We address this challenge by proposing a novel approach for mislabeled data detection that leverages training dynamics. Our method is grounded in the key observation that correctly labeled samples exhibit consistent entropy decrease during training, while mislabeled samples maintain relatively high entropy throughout the training process. Building on this insight, we introduce a signed entropy integral (SEI) statistic that captures both the magnitude and temporal trend of prediction entropy across training epochs. SEI is broadly applicable to classification networks and demonstrates particular effectiveness when integrated with contrastive language-image pretraining (CLIP) architectures. Through extensive experiments on four medical imaging datasets -- a domain particularly susceptible to labeling errors due to diagnostic complexity -- spanning diverse modalities and pathologies, we demonstrate that SEI achieves state-of-the-art performance in mislabeled data identification, outperforming existing methods while maintaining computational efficiency and implementation simplicity. Our code is available at https://github.com/MedAITech/SEI.

cs.CV

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.

cs.AI

Helical Rashba-exchange gauge field drives a uniaxial pair density wave in EuRbFe$_4$As$_4$

The recent discovery of an intrinsic, zero-field pair density wave (PDW) in the iron-pnictide superconductor EuRbFe$_4$As$_4$ poses a fundamental puzzle: how does a unidirectional, nanometer-scale superconducting modulation arise spontaneously below the magnetic ordering temperature? Here we show that the interplay of Rashba spin-orbit coupling -- induced by the locally non-centrosymmetric FeAs layers -- and the period-four helical Eu$^{2+}$ exchange field generates a layer-rotating effective $U(1)$ gauge field for the Cooper pairs. Because this gauge field shares the symmetry of the Fe $3d_{xz}/3d_{yz}$ orbital doublet, it drives an orbital-selective, finite-momentum pairing instability. Using a Ginzburg-Landau theory on the magnetic unit cell, we demonstrate that this mechanism naturally stabilizes a strictly uniaxial Bloch superconducting state at the experimentally observed wavelength, accompanied by spontaneous interlayer loop currents accessible to muon-spin relaxation or scanning SQUID microscopy.

cond-mat.supr-con

Parametrically driven pure-quartic solitons

Parametrically driven solitons are self-trapped modes in various physical settings, including optics, magnetics, etc. So far, the analysis was focused on the existence, stability, and dynamics of such solitons in systems including the second-order group-velocity dispersion (GVD), linear loss, parametric gain, and cubic nonlinearity. Here, we report the existence of quiescent parametrically driven pure-quartic solitons (PDPQSs) in the full system, and moving PDPQSs in the absence of losses. A systematic analysis reveals stability domains for the solitons in the system's parameter space. Evolution of unstable states is explored too, and it is demonstrated that collisions between traveling stable PDPQSs are elastic.

physics.optics

In-Context Positive-Unlabeled Learning

Positive-unlabeled (PU) learning addresses binary classification when only a set of labeled positives is available alongside a pool of unlabeled samples drawn from a mixture of positives and negatives. Existing PU methods typically require dataset-specific training or iterative optimization, which limits their applicability when many tasks must be solved quickly or with little tuning. We introduce PUICL, a pretrained transformer that solves PU classification entirely through in-context learning. PUICL is pretrained on synthetic PU datasets generated from randomly instantiated structural causal models, exposing it to a wide range of feature-label relationships and class-prior configurations. At inference time, PUICL receives the labeled positives and the unlabeled samples as a single input and returns class probabilities for the unlabeled rows in one forward pass, with no gradient updates or per-task fitting. On 20 semi-synthetic PU benchmarks derived from the UCI Machine Learning Repository, OpenML, and scikit-learn, PUICL outperforms four standard PU learning baselines in average AUC and accuracy, and is competitive on F1-score. These results show that the in-context learning paradigm extends naturally beyond fully supervised tabular prediction to the semi-supervised PU setting.

stat.ML

A semiparametric two-sample homogeneity test with nonignorable nonresponse using callback data

Testing the homogeneity of two distributions is fundamental in statistics, but classical procedures may fail under nonignorable nonresponse. In many surveys, callback data record repeated contact attempts and provide auxiliary information about the response mechanism. We develop a semiparametric framework for two-sample homogeneity testing that explicitly incorporates such information. The response mechanism is modeled by a flexible semiparametric callback model, while the two population distributions are linked through a density ratio model. Within this unified framework, we propose an empirical likelihood ratio test for distributional homogeneity and show that, under the null hypothesis, it has a Wilks-type chi-square limit. To facilitate computation, we develop an efficient expectation-maximization-type algorithm. Simulation results show that the proposed method controls type I error well and achieves substantially higher power than existing methods that ignore nonignorable missingness. An application to real survey income data illustrates its practical value.

stat.ME

VistaBot: View-Robust Robot Manipulation via Spatiotemporal-Aware View Synthesis

Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that integrates feed-forward geometric models with video diffusion models to achieve view-robust closed-loop manipulation without requiring camera calibration at test time. Our approach consists of three key components: 4D geometry estimation, view synthesis latent extraction, and latent action learning. VistaBot is integrated into both action-chunking (ACT) and diffusion-based ($\pi_0$) policies and evaluated across simulation and real-world tasks. We further introduce the View Generalization Score (VGS) as a new metric for comprehensive evaluation of cross-view generalization. Results show that VistaBot improves VGS by 2.79$\times$ and 2.63$\times$ over ACT and $\pi_0$, respectively, while also achieving high-quality novel view synthesis. Our contributions include a geometry-aware synthesis model, a latent action planner, a new benchmark metric, and extensive validation across diverse environments. The code and models will be made publicly available.

cs.RO

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance

Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level knowledge and spatial awareness. To address these challenges, we propose PokeVLA, a lightweight yet powerful foundation model for embodied manipulation that effectively infuses vision-language understanding into action learning. Our framework introduces a two-stage training paradigm: first, we pre-train a compact vision-language model (PokeVLM) on a curated multimodal dataset of 2.4M samples encompassing spatial grounding, affordance, and embodied reasoning tasks; second, we inject manipulation-relevant representations into the action space through multi-view goal-aware semantics learning, geometry alignment, and a novel action expert. Extensive experiments demonstrate state-of-the-art performance on the LIBERO-Plus benchmark and in real-world deployment, outperforming comparable baselines in success rate and robustness under diverse perturbations. To foster reproducibility and community progress, we will open-source our code, model weights, and the scripts for the curated pre-training dataset. Project page: https://getterupper.github.io/PokeVLA

cs.RO