SearcharxivSearch

arXiv subjects

Huiming Chen

Publications and source records attributed to Huiming Chen.

7 recordsLinked to original sources

Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. For the evaluated Bridge/WidowX and physical pick-place deployments, an explicit execution contract assigns PICK to source-conditioned OFG and PLACE to Pointing, providing direct stage-aligned spatial targets. Pointing-VLA achieves SOTA performance on Bridge/WidowX, averaging 72.9\% across the evaluated four-task set without Bridge-specific finetuning under collision-enabled CuRobo execution. Pointing and OFG show complementary strengths across native and cross-dataset evaluations. The OFG/contact readout transfers to NORA-1.5, preserving or improving success while reducing recorded controller time by more than 20$\times$; typed heads are also 6.68--6.90$\times$ faster than Embodied-R1 text decoding on a shared external suite. When integrated as spatial guidance for a $\pi_{0.5}$ action policy, Pointing-VLA raises autonomous real-robot success from 52.7\% to 80.7\% across three visual contexts. These results establish typed spatial readouts as an efficient, inspectable interface between embodied reasoning and robot execution.

cs.RO

Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enters the executed prompt. In a matched audit, raw appended text reduces mean success from 92.47\% to 3.00\%, while meaningful and length-matched meaningless appends both fail on all 500 states. This result identifies \emph{prompt-form collapse}: changing the instruction form, rather than adding useful semantics, can dominate execution. We introduce TOWN-VLA (Think Only When Needed), a prompt-authority interface that separates candidate generation from permission to alter the policy input. A fixed compatibility rule authorizes a canonical compact instruction; otherwise, the interface restores the original Base prompt exactly. Across 900 audited routes, every route follows this contract: 525 routes recover Base with matching hashes, and all 375 authorized prompts preserve the task signature. On a matched $4\times7$ LIBERO-Plus evaluation with 10{,}030 episodes per method, success rises from 69.5\% to 73.1\% ($+362$ episodes; 95\% CI 1.89--5.45 points), improving on six perturbation axes and all four suites. On a physical PiPER arm with a frozen \pizerofive{} checkpoint, success rises from 52.7\% to 78.7\% over 150 trials per method ($p=3.16\times10^{-6}$). Prompt authority is enforceable for a frozen controller; oracle-free admission calibration is the next deployment target.

cs.RO

Advancements in Federated Learning: Models, Methods, and Privacy

Federated learning (FL) is a promising technique for addressing the rising privacy and security issues. Its main ingredient is to cooperatively learn the model among the distributed clients without uploading any sensitive data. In this paper, we conducted a thorough review of the related works, following the development context and deeply mining the key technologies behind FL from both theoretical and practical perspectives. Specifically, we first classify the existing works in FL architecture based on the network topology of FL systems with detailed analysis and summarization. Next, we abstract the current application problems, summarize the general techniques and frame the application problems into the general paradigm of FL base models. Moreover, we provide our proposed solutions for model training via FL. We have summarized and analyzed the existing FedOpt algorithms, and deeply revealed the algorithmic development principles of many first-order algorithms in depth, proposing a more generalized algorithm design framework. Based on these frameworks, we have instantiated FedOpt algorithms. As privacy and security is the fundamental requirement in FL, we provide the existing attack scenarios and the defense methods. To the best of our knowledge, we are among the first tier to review the theoretical methodology and propose our strategies since there are very few works surveying the theoretical approaches. Our survey targets motivating the development of high-performance, privacy-preserving, and secure methods to integrate FL into real-world applications.

cs.AI

LoSAC: An Efficient Local Stochastic Average Control Method for Federated Optimization

Federated optimization (FedOpt), which targets at collaboratively training a learning model across a large number of distributed clients, is vital for federated learning. The primary concerns in FedOpt can be attributed to the model divergence and communication efficiency, which significantly affect the performance. In this paper, we propose a new method, i.e., LoSAC, to learn from heterogeneous distributed data more efficiently. Its key algorithmic insight is to locally update the estimate for the global full gradient after {each} regular local model update. Thus, LoSAC can keep clients' information refreshed in a more compact way. In particular, we have studied the convergence result for LoSAC. Besides, the bonus of LoSAC is the ability to defend the information leakage from the recent technique Deep Leakage Gradients (DLG). Finally, experiments have verified the superiority of LoSAC comparing with state-of-the-art FedOpt algorithms. Specifically, LoSAC significantly improves communication efficiency by more than $100\%$ on average, mitigates the model divergence problem and equips with the defense ability against DLG.

cs.DC

Joint Channel Estimation and Training Signal Design for Two-way MIMO Relay Systems

In this paper, a two-stage channel estimation scheme for two-way MIMO relay systems with a single relay antenna is proposed. The backward channel is estimated by using linear minimum mean square estimator (LMMSE) at the first stage, where the optimal training signal is designed. We then mainly focus on the forward channel estimation by using singular value decomposition (SVD) based maximum likelihood method, and the related training signal is proposed. We note that the forward channel estimator is nonlinear and by analyzing the asymptotic Bayesian Cramer-rao Lower Bound (BCRLB), we seek BCRLB as the criterion for training signal design. Finally, the numerical results show that the proposed training signal can improve the MSE performance.

cs.IT

Training Design and Two-stage Channel Estimation for Correlated Two-way MIMO Relay Systems

This paper addresses the training signal design for the channel estimation in two-way multiple-input-and-multipleoutput (MIMO) relay systems, where the channels are correlated. We first derive the backward channel estimator with the optimal training signal sent by the relay node. Given the estimated backward channels and the probabilistic knowledge of the estimation error, we mainly focus on the forward channel estimation and the related training signal design. We further propose a novel training signal. The design criterion is to minimize the relaxation of the total mean square error (MSE) of the forward channel estimators, which is conditioned on the estimated backward channels. Finally, the numerical results show that the proposed training signal can improve the MSE performance.

cs.IT

Quantum pumping with adiabatically modulated barriers in graphene

We study the adiabatic quantum pumping characteristics in the graphene modulated by two oscillating gate potentials out of phase. The angular and energy dependence of the pumped current is presented. The direction of the pumped current can be reversed when a high barrier demonstrates stronger transparency than a low one, which results from the Klein paradox. The underlying physics of the pumping process is illuminated.

cond-mat.mes-hall