SearcharxivSearch

arXiv subjects

Min Wang

Publications and source records attributed to Min Wang.

At least 19 recordsLinked to original sources

Perceptually Regularized Diffusion Model for Image Super-Resolution

Image super-resolution, which aims to reconstruct high-resolution images from their low-resolution observations, is fundamental to medical imaging, remote sensing, surveillance, microscopy, and scientific visualization. Traditional model-based methods formulate super-resolution as an inverse problem with hand-crafted regularization priors. While interpretable and theoretically grounded, they rely on fixed assumptions and require computationally intensive iterative solvers. Deep learning methods offer data-driven flexibility by learning nonlinear mappings from low- to high-resolution images, among which diffusion models have achieved particularly impressive perceptual quality. However, the standard diffusion training objective is a pixel-domain noise-prediction loss that does not explicitly enforce perceptual fidelity, which can lead to oversmoothing and loss of fine image structure. To address these limitations, we propose a perceptually regularized diffusion framework that incorporates prior knowledge through perceptual-loss-based regularization, improving training convergence and encouraging the recovery of meaningful image features. Experiments on benchmark datasets demonstrate improved perceptual quality and competitive distortion metrics, highlighting the effectiveness of regularization for diffusion-based super resolution.

eess.IV

ParaTempo: Efficient Parallel Reasoning via Temporal Confidence

Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managing these parallel paths typically rely on final-answer consensus, local token confidence, or isolated intermediate probes. However, these signals are often delayed, weakly tied to actual reasoning progress, or too noisy for dynamic, branch-level control. To address these limitations, we introduce ParaTempo, a training-free asynchronous parallel reasoning framework. ParaTempo is driven by temporal confidence, a branch-local measure of answer-space convergence. Each branch is periodically probed for a tentative answer probability distribution, and temporal confidence quantifies how sharply the recent intermediate probes concentrate on a dominant answer. Once sufficient evidence has accumulated, ParaTempo drives its entire control process from this single signal: low-confidence branches are pruned, branches that persistently commit to their dominant answer are retired early, freed computation is reallocated by forking new branches, and generation stops globally once the confidence-weighted vote concentrates. Without requiring synchronization among reasoning trajectories, ParaTempo adaptively allocates computation based on branch-level convergence. Experiments on challenging mathematical and scientific reasoning benchmarks show that ParaTempo reduces average latency by 21.8-32.2% and total token usage by 18.1-30.3% while maintaining competitive accuracy. Moreover, temporal confidence exhibits stronger temporal stability and predictive power for future branch convergence than token-level and instantaneous signals.

cs.AI

Asymmetric fractional coupled magnetizable piezoelectric beams with infinite viscoelastic memory: polynomial decay and sharpness of the decay rate

This paper studies the long-time dynamics of a coupled hyperbolic system for magnetizable piezoelectric beams with infinite viscoelastic memory. Memory dissipation is characterized by $A^\alpha$, the fractional power of a positive self-adjoint operator $A$ with $\alpha\in[0,1)$. An asymmetric fractional magnetoelectric coupling is adopted: the mechanical-to-magnetic coupling uses integer-order operator $A$, while the magnetic feedback to mechanics is governed by fractional operator $A^\beta$ ($\beta\in[0,1)$). This model bridges the gap between the well-studied integer coupling case ($\beta=1$) and the unsolved fully fractional symmetric coupling problem, offering a universal framework for related coupled systems. Under mild assumptions on memory kernels and system parameters, we prove well-posedness via semigroup theory and derive an explicit polynomial decay estimate for smooth initial data: \[ \|X(t)\|_{\mathcal H} \le C t^{-\frac{1}{4-2\beta-2\alpha}}\|X_0\|_{D(\mathcal A)},\quad \forall\,t\ge 1, \] where the decay exponent is explicitly determined by $\alpha$ and $\beta$. For exponentially decaying memory kernels, the decay rate is sharp if stiffness coefficients satisfy $\alpha_1\ne\alpha_2$. When $\alpha_1=\alpha_2$, we only obtain an upper bound $\delta\le \frac{1}{3-\beta-2\alpha}$ for decay index $\delta$, leaving the optimal rate open. Comparisons with integer feedback coupling ($\beta=1$) show fractional feedback ($\beta<1$) slows energy decay. It demonstrates that $\beta$ weakens indirect damping and degrades structural stabilization. The results reveal the intrinsic interaction between fractional memory dissipation and fractional coupling in strongly coupled dissipative systems.

math.AP

Multi-Angular Reflectance Anisotropy Observed from UAV Multispectral Imagery

UAV multispectral imagery naturally contains multi-angular observations due to low flight altitude and wide field-of-view imaging, which may introduce geometry-driven radiometric variability. This study proposes a geometry-aware multi-angular observation extraction workflow to quantify observation-geometry effects from a BRDF perspective. Specifically, camera intrinsics and extrinsics are refined via structure-from-motion (SFM), and homogeneous regions annotated on an orthomosaic are reprojected onto multiple raw sub-images acquired from different viewpoints. This enables joint extraction of multi-band reflectance and observation geometry parameters for the same ground targets under varying viewing directions. The extracted observations are further analyzed using band-wise polar visualization in the (VZA, RAA) domain. Results on a grassland target show clear reflectance anisotropy across ten bands, with red-edge and nearinfrared bands exhibiting 119-137% variability between maximum and minimum reflectance, indicating non-negligible observation-geometry effects on radiometric consistency.

cs.CV

Scalable native signed optical computing enabled by dual-wavelength incoherent multiplexing

Incoherent photonic neural networks (PNNs) provide a robust platform for analog optical computing, yet efficient implementation of native signed operations remains challenging. Existing incoherent PNNs approaches often require additional spatial channels or temporal encoding steps to represent bipolar input signals, resulting in hardware overhead that scales with system size. Here, we demonstrate a dual-wavelength incoherent photonic architecture that natively supports both signed inputs and signed weights on a thin-film lithium niobate platform. By encoding complementary signal components onto two wavelength channels and performing computation within a shared physical path, the proposed scheme eliminates duplicated weighting units. As a result, the additional hardware overhead associated with signed computation remains constant per multiply accumulate operation, independent of matrix size. The fabricated device exhibits a modulation bandwidth exceeding 40 GHz and achieves four-quadrant optical multiplication with a standard deviation error of 1.27%. System-level functionality is validated through neural-network classification, achieving 95.1% accuracy on the Moons dataset and 91.63% on MNIST. These results establish a practical route toward scalable incoherent photonic computing systems with native bipolar processing capability.

physics.optics

Hybrid-Integrated DFB-Laser-Coupled 1 * 8 Thin-Film Lithium Niobate Modulator Array for High-Speed Parallel Optical Transmitters

Thin-film lithium niobate (TFLN) electro-optic modulators are attractive for high-speed optical interconnects, but scalable transmitter architectures require not only high modulation bandwidth but also multi-channel optical power distribution and practical laser-to-chip integration. Here, we demonstrate a hybrid-integrated 1 * 8 TFLN electro-optic modulator array passively butt-coupled to a 1550 nm distributed-feedback laser. The chip integrates a three-stage cascaded 1 * 2 multimode-interference splitter, spot-size converters, eight traveling-wave Mach-Zehnder modulators, thermal tuning electrodes, and on-chip 50 {\Omega} terminations. The cascaded splitter provides uniform optical power distribution with a maximum normalized power deviation of 9.7%, while the optimized electrodes enable electro-optic 3 dB bandwidths exceeding 40 GHz for all channels. The measured half-wave voltages are 3.60-3.83 V, corresponding to V{\pi}L products of 2.52-2.68 V cm for a 7 mm modulation length, and the extinction ratio reaches approximately 25 dB. The bare-chip insertion loss is 15.19-16.55 dB, and DFB laser bonding introduces an additional coupling loss of approximately 5 dB while preserving channel uniformity. These results establish a practical TFLN-based multi-channel modulator platform and represent a step toward compact hybrid-integrated optical transmitters for high-speed parallel interconnects.

physics.optics

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models

Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervision for execution drift, while failed rollouts are often discarded. We introduce RePO-VLA, a recovery-driven policy optimization framework that assigns distinct roles to success, recovery, and failure trajectories. RePO-VLA first applies Recovery-Aware Initialization (RAI), slicing recovery segments and resetting history so corrective actions depend on the current adverse state rather than the preceding failure. It then learns a Progress-Aware Semantic Value Function (PAS-VF), aligning spatiotemporal trajectory features with instructions and successful references. The resulting labels salvage useful failure prefixes via reliability decay, while low-value labels mark drift and terminal breakdowns, teaching differences among nominal, failed, and corrective actions. The data engine turns adverse states into planner-generated or human-collected corrective rollouts, teaching recovery to the success manifold. Value-Conditioned Refinement (VCR) trains the policy to prefer high-progress actions. At deployment, a fixed high value ($v=1.0$) biases actions toward the learned success manifold without online failure detectors or heuristic retries. We introduce FRBench, with standardized error injection and recovery-focused evaluation. Across simulated and real-world bimanual tasks, RePO-VLA improves robustness, raising adversarial success from 20% to 75% on average and up to 80% in scaled real-world trials.

cs.RO

A Multimodal Pre-trained Network for Integrated EEG-Video Seizure Detection

Reliable seizure detection in mouse models is essential for preclinical epilepsy research, yet manual review of synchronized video-EEG recordings is labor-intensive and single-modality systems fail for complementary reasons: video-based methods are easily confounded by benign behaviors, whereas EEG-based methods are vulnerable to ictal motion artifacts. We present EEGVFusion, a multimodal framework that combines self-supervised EEG representation learning, spatio-temporal video encoding, optimal-transport alignment, and bidirectional cross-attention to integrate neural and behavioral evidence. We also curate an expert-annotated dataset of synchronized EEG and video recordings comprising 93 sessions from 15 mice for training and evaluation. In the random-session split, EEGVFusion achieved a Balanced Accuracy of 0.9957 with perfect event sensitivity and an Event FAR of 0.6250 FP/h, indicating strong seizure detection performance with a low false-alarm burden. In a single held-out-subject evaluation with Subject 110 reserved for testing, EEGVFusion achieved a Balanced Accuracy of 0.9718 and reduced Event FAR from 2.7250 FP/h for the EEG-only counterpart to 0.4833 FP/h while preserving perfect event sensitivity. Targeted ablations further showed that EEG pre-training and OT alignment help reduce false alarms while preserving event sensitivity.

cs.CV

A B-Spline Function Based 3D Point Cloud Unwrapping Scheme for 3D Fingerprint Recognition and Identification

Three-dimensional (3D) fingerprint recognition and identification offer several advantages over traditional two-dimensional (2D) recognition systems. The contactless nature of 3D fingerprints enhances hygiene and security, reducing the risk of contamination and spoofing. In addition to surface ridge and valley patterns, 3D fingerprints capture depth, curvature, and shape information, enabling the development of more precise and robust authentication systems. Despite recent advancements, significant challenges remain. The topological height of fingerprint pixels complicates the extraction of ridge and valley patterns. Furthermore, registration issues limit the acquisition process, requiring consistent direction and orientation across all samples. To address these challenges, this paper introduces a method that unwraps 3D fingerprints, represented as 3D point clouds, using B-spline curve fitting to mitigate height variation and reduce registration limitations. The unwrapped point cloud is then converted into a grayscale image by mapping the relative heights of the points. This grayscale image is subsequently used for recognition through conventional 2D fingerprint identification methods. The proposed approach demonstrated superior performance in 3D fingerprint recognition, achieving Equal Error Rates (EERs) of 0.2072%, 0.26%, and 0.22% across three experiments, outperforming existing methods. Additionally, the method surpassed 3D fingerprint flattening technique in both recognition and identification during cross-session experiments, achieving an EER of 1.50% when fingerprints with varying registrations were included.

cs.CV

AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark), a long-horizon benchmark for hidden-target identification over galleries of visually similar images. In AMIGO, the oracle privately selects a target image, and the model must recover it by asking a sequence of attribute-focused Yes/No/Unsure questions under a strict protocol that penalizes invalid actions with Skip. This setting stresses (i) question selection under uncertainty, (ii) consistent constraint tracking across turns, and (iii) fine-grained discrimination as evidence accumulates. AMIGO also supports controlled oracle imperfections to probe robustness and verification behavior under inconsistent feedback. We instantiate AMIGO with Guess My Preferred Dress task and report metrics covering both outcomes and interaction quality, including identification success, evidence verification, efficiency, protocol compliance, noise tolerance, and trajectory-level diagnostics.

cs.LG

Highly-efficient, narrow-linewidth Brillouin microlasers implemented in compact thin-film lithium niobate microresonators

Stimulated Brillouin microlasers offer chip-scale light sources with high spectral purity and low phase noise--key attributes for applications spanning precision metrology, quantum technologies, and coherent information processing. However, simultaneously bringing both pump and scattered waves into resonance often compromises photon confinement or modal volume, resulting in limited conversion efficiency and elevated thresholds. In this work, a novel approach is proposed to generate Brillouin microlasers with high efficiency, low threshold, and narrow linewidth, by combining a cross-polarized stimulated Brillouin scattering scheme with intentional Stokes mode splitting to compensate for mode detuning. Triple-resonance and phase-matching conditions are simultaneously achieved in a 114-um-diameter thin-film lithium niobate (TFLN) microresonator, enabling precise alignment with both the ~10-GHz Brillouin shift and the ~100-MHz narrow gain bandwidth. The resulting Brillouin microlaser achieves a narrow intrinsic linewidth of 2.88 Hz, a short-term integral linewidth of 185 Hz, an on-chip conversion efficiency of 57.92%, and a pump threshold as low as 1.03 mW. Both the conversion efficiency and the lasing threshold represent record-high performance for the TFLN platform to date.

physics.optics

Structural Action Transformer for 3D Dexterous Manipulation

Achieving human-level dexterity in robots via imitation learning from heterogeneous datasets is hindered by the challenge of cross-embodiment skill transfer, particularly for high-DoF robotic hands. Existing methods, often relying on 2D observations and temporal-centric action representation, struggle to capture 3D spatial relations and fail to handle embodiment heterogeneity. This paper proposes the Structural Action Transformer (SAT), a new 3D dexterous manipulation policy that challenges this paradigm by introducing a structural-centric perspective. We reframe each action chunk not as a temporal sequence, but as a variable-length, unordered sequence of joint-wise trajectories. This structural formulation allows a Transformer to natively handle heterogeneous embodiments, treating the joint count as a variable sequence length. To encode structural priors and resolve ambiguity, we introduce an Embodied Joint Codebook that embeds each joint's functional role and kinematic properties. Our model learns to generate these trajectories from 3D point clouds via a continuous-time flow matching objective. We validate our approach by pre-training on large-scale heterogeneous datasets and fine-tuning on simulation and real-world dexterous manipulation tasks. Our method consistently outperforms all baselines, demonstrating superior sample efficiency and effective cross-embodiment skill transfer. This structural-centric representation offers a new path toward scaling policies for high-DoF, heterogeneous manipulators.

cs.RO

CosmicWeb-21cm array: A New Radio Observation Array Design for 21cm Cosmology

This paper presents the CosmicWeb-21cm array, a novel radio interferometer designed to overcome the key challenges in 21 cm cosmology. Its core innovations include: (1) a multi-scale nested geometry combining a hexagonal core with logarithmic spiral arms for excellent UV coverage and calibration robustness; (2) an intelligent non-uniform frequency sampling strategy that adapts resolution to foreground and signal characteristics, reducing data volume while preserving information; and (3) a machine-learning-enhanced, physics-informed processing pipeline that achieves 99.7\% foreground removal efficiency; (4) a dual-polarization crossed dipole integrated with a dielectric lens and cryogenically cooled LNA, achieving stable beam patterns and low noise temperature ($<35$ K) across 50-250 MHz. These co-designed advances enable high sensitivity mapping of the Epoch of Reionization, dark energy constraints and cosmic-web structure.

astro-ph.IM

Primary-Fine Decoupling for Action Generation in Robotic Imitation

Multi-modal distribution in robotic manipulation action sequences poses critical challenges for imitation learning. To this end, existing approaches often model the action space as either a discrete set of tokens or a continuous, latent-variable distribution. However, both approaches present trade-offs: some methods discretize actions into tokens and therefore lose fine-grained action variations, while others generate continuous actions in a single stage tend to produce unstable mode transitions. To address these limitations, we propose Primary-Fine Decoupling for Action Generation (PF-DAG), a two-stage framework that decouples coarse action consistency from fine-grained variations. First, we compress action chunks into a small set of discrete modes, enabling a lightweight policy to select consistent coarse modes and avoid mode bouncing. Second, a mode conditioned MeanFlow policy is learned to generate high-fidelity continuous actions. Theoretically, we prove PF-DAG's two-stage design achieves a strictly lower MSE bound than single-stage generative policies. Empirically, PF-DAG outperforms state-of-the-art baselines across 56 tasks from Adroit, DexArt, and MetaWorld benchmarks. It further generalizes to real-world tactile dexterous manipulation tasks. Our work demonstrates that explicit mode-level decoupling enables both robust multi-modal modeling and reactive closed-loop control for robotic manipulation.

cs.RO

A Real-World Grasping-in-Clutter Performance Evaluation Benchmark for Robotic Food Waste Sorting

Food waste management is critical for sustainability, yet inorganic contaminants hinder recycling potential. Robotic automation accelerates sorting through automated contaminant removal. Nevertheless, the diverse and unpredictable nature of contaminants introduces major challenges for reliable robotic grasping. Grasp performance benchmarking provides a rigorous methodology for evaluating these challenges in underexplored field contexts like food waste sorting. However, existing approaches suffer from limited simulation datasets, over-reliance on simplistic metrics like success rate, inability to account for object-related pre-grasp conditions, and lack of comprehensive failure analysis. To address these gaps, this work introduces GRAB, a real-world grasping-in-clutter (GIC) performance benchmark incorporating: (1) diverse deformable object datasets, (2) advanced 6D grasp pose estimation, and (3) explicit evaluation of pre-grasp conditions through graspability metrics. The benchmark compares industrial grasping across three gripper modalities through 1,750 grasp attempts across four randomized clutter levels. Results reveal a clear hierarchy among graspability parameters, with object quality emerging as the dominant factor governing grasp performance across modalities. Failure mode analysis shows that physical interaction constraints, rather than perception or control limitations, constitute the primary source of grasp failures in cluttered environments. By enabling identification of dominant factors influencing grasp performance, GRAB provides a principled foundation for designing robust, adaptive grasping systems for complex, cluttered food waste sorting.

cs.RO

Solving and learning advective multiscale Darcian dynamics with the Neural Basis Method

Physics-governed models are increasingly paired with machine learning for accelerated predictions, yet most "physics--informed" formulations treat the governing equations as a penalty loss whose scale and meaning are set by heuristic balancing. This blurs operator structure, thereby confounding solution approximation error with governing-equation enforcement error and making the solving and learning progress hard to interpret and control. Here we introduce the Neural Basis Method, a projection-based formulation that couples a predefined, physics-conforming neural basis space with an operator-induced residual metric to obtain a well-conditioned deterministic minimization. Stability and reliability then hinge on this metric: the residual is not merely an optimization objective but a computable certificate tied to approximation and enforcement, remaining stable under basis enrichment and yielding reduced coordinates that are learnable across parametric instances. We use advective multiscale Darcian dynamics as a concrete demonstration of this broader point. Our method produce accurate and robust solutions in single solves and enable fast and effective parametric inference with operator learning.

math.NA

Parametrization of subgrid scales in long-term simulations of the shallow-water equations using machine learning and convex limiting

We present a method for parametrizing sub-grid processes in the Shallow Water equations. We define coarse variables and local spatial averages and use a feed-forward neural network to learn sub-grid fluxes. Our method results in a local parametrization that uses a four-point computational stencil, which has several advantages over globally coupled parametrizations. We demonstrate numerically that our method improves energy balance in long-term turbulent simulations and also accurately reproduces individual solutions. The long-term simulations refer to numerical studies where a fluid flow is simulated over a duration long enough to reach a statistical steady state. The neural network parametrization can be easily combined with flux limiting to reduce oscillations near shocks. More importantly, our method provides reliable parametrizations, even in dynamical regimes that are not included in the training data.

physics.flu-dyn

Offline Meta-Reinforcement Learning with Flow-Based Task Inference and Adaptive Correction of Feature Overgeneralization

Offline meta-reinforcement learning (OMRL) combines the strengths of learning from diverse datasets in offline RL with the adaptability to new tasks of meta-RL, promising safe and efficient knowledge acquisition by RL agents. However, OMRL still suffers extrapolation errors due to out-of-distribution (OOD) actions, compromised by broad task distributions and Markov Decision Process (MDP) ambiguity in meta-RL setups. Existing research indicates that the generalization of the $Q$ network affects the extrapolation error in offline RL. This paper investigates this relationship by decomposing the $Q$ value into feature and weight components, observing that while decomposition enhances adaptability and convergence in the case of high-quality data, it often leads to policy degeneration or collapse in complex tasks. We observe that decomposed $Q$ values introduce a large estimation bias when the feature encounters OOD samples, a phenomenon we term ''feature overgeneralization''. To address this issue, we propose FLORA, which identifies OOD samples by modeling feature distributions and estimating their uncertainties. FLORA integrates a return feedback mechanism to adaptively adjust feature components. Furthermore, to learn precise task representations, FLORA explicitly models the complex task distribution using a chain of invertible transformations. We theoretically and empirically demonstrate that FLORA achieves rapid adaptation and meta-policy improvement compared to baselines across various environments.

cs.LG