SearcharxivSearch

arXiv subjects

Bin Zheng

Publications and source records attributed to Bin Zheng.

At least 19 recordsLinked to original sources

Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG

Graph-based retrieval-augmented generation (RAG) provides a scalable paradigm for long-video understanding, but existing systems typically inherit a fixed temporal granularity from video segmentation when constructing their retrieval index. We argue that this design unnecessarily couples indexing granularity with evidence granularity: coarse representations can often suffice for locating relevant temporal regions, while fine-grained evidence remains important for downstream reasoning. We propose \textbf{Density-Aware Graph Construction (DAGC)}, a training-free approach that decouples a query-independent coarse retrieval index from the original fine-grained evidence space. DAGC constructs a compact, density-adaptive graph index by merging visually redundant neighboring chunks, while preserving mappings to the original temporal units. Retrieved coarse regions are subsequently expanded back to the original chunk granularity for fine-grained evidence refinement and answer generation. Experiments on MLVU, VideoMME, and LongVideoBench show that DAGC retains only about 40--50\% of the original graph nodes and achieves $1.3$--$1.7\times$ end-to-end wall-clock acceleration while preserving approximately 99\% of the original QA performance. The gains transfer across different LVLM backbones and video RAG pipelines, suggesting that long-video RAG need not maintain the same temporal granularity for indexing and evidence reasoning.

cs.CV

Non-uniform swelling of polyelectrolyte hydrogels: effects of charge regulation

We investigate the impact of charge regulation (CR) on the non-uniform swelling behavior of polyelectrolyte hydrogels. The Poisson-Boltzmann theory with electro-elastic coupling between the local polymer density and elastic deformation is considered. We investigate the spatial distributions of the elastic displacement and polymer density under different salt concentrations and compare charge-regulated gels with fixed-charge (non-CR) gels of the same net charge. Our results show that the CR induces spatially varying charge fractions, which strengthen the electro-elastic response and lead to stronger non-uniform swelling compared with non-CR gels. These findings provide a theoretical basis for understanding and controlling non-uniform swelling in responsive polyelectrolyte hydrogels.

cond-mat.soft

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

Human-in-the-loop reinforcement learning (HIL-RL) enables robots to learn contact-rich manipulation from limited real-world interaction, but deployment exposes three coupled limitations: static visual reward models fail under scene changes; independently sampled actions cause temporally inconsistent motion; and vision-based policies remain sensitive to appearance shifts. We present EvoHIL, a unified framework that adapts the reward model, action generator, and visual do main within a staged human-in-the-loop learning process. First, self-evolving reward (SER) adapts the success classifier from human-confirmed positives and provisional weak negatives. Second, Action Flow Stabilization (AFS) generates temporally coherent action chunks through flow matching, grounding policy updates in executed action prefixes and demonstrated behavior. Third, retention-aware offline fine-tuning replays relit interaction data while anchoring the AFS actor-critic to prior behavior, adapting the visual domain without additional robot interaction. Across six manipulation tasks on Franka FR3 and SO-101 arms under a controlled lighting shift, EvoHIL improves task success, agreement with human-confirmation labels, motion smoothness, and completion time relative to human-in-the-loop and imitation baselines.Project page: https://anonymous4366.github.io/EvoHIL/

cs.RO

Covariant Onsager and Onsager-Machlup principles for active and inertial dynamics

The Onsager principle provides a variational route to the phenomenological equations of dissipative dynamics through the minimization of the Rayleighian. We develop a covariant formulation of the Onsager principle for active and inertial systems, ensuring geometric consistency under coordinate transformations. To further incorporate thermal fluctuations, we formulate the Onsager-Machlup principle for active and inertial systems by considering the Onsager-Machlup functional and the corresponding path probability for stochastic trajectories. Requiring that the path probability obeys the detailed fluctuation theorem, we show that the extended Onsager-Machlup theory is consistent with stochastic thermodynamics. The extended OP and OMP offer a unified and useful variational framework for deriving the dynamical equations of active and inertial systems.

cond-mat.soft

PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads

Accurately forecasting GPU workloads is essential for AI infrastructure, enabling efficient scheduling, resource allocation, and power management. Modern workloads are highly volatile, multiple periodicity, and heterogeneous, making them challenging for traditional predictors. We propose PRISM, a primitive-based compositional forecasting framework combining dictionary-driven temporal decomposition with adaptive spectral refinement. This dual representation extracts stable, interpretable workload signatures across diverse GPU jobs. Evaluated on large-scale production traces, PRISM achieves state-of-the-art results. It significantly reduces burst-phase errors, providing a robust, architecture-aware foundation for dynamic resource management in GPU-powered AI platforms.

cs.DC

Matrix-Free Stabilized BDF Schemes for Semilinear Parabolic Equations with Unconditional Maximum Bound Principle Preservation and Energy Stability

We develop a family of stabilized backward differentiation formula (sBDF) schemes of orders one through four for semilinear parabolic equations. The proposed methods are designed to achieve three properties that are rarely available simultaneously in high-order time discretizations: unconditional preservation of the maximum bound principle (MBP), unconditional discrete energy stability, and practical matrix-free implementation. The construction integrates carefully designed stabilization terms, fixed-point iterations, and a pointwise cut-off strategy. The nonlinear algebraic systems arising from the implicit sBDF discretizations are solved by fixed-point iteration, resulting in fully matrix-free algorithms. This makes the approach particularly attractive for practical computations on general domains and under mixed boundary conditions, where FFT-based exponential time differencing methods are often unavailable or inefficient. We further present a unified analysis for the fully implemented schemes, explicitly incorporating the interplay among time discretization, nonlinear iteration, and cut-off. Unconditional contractivity of the fixed-point iterations and error estimates are established. For the Allen-Cahn equation, we additionally prove an unconditional discrete energy dissipation law. Numerical experiments confirm the theoretical convergence rates and demonstrate the robustness and efficiency of the proposed methods, particularly relative to ETD-based approaches for problems with mixed boundary conditions.

math.NA

Fronthaul-Efficient Distributed Cooperative 3D Positioning with Quantized Latent CSI Embeddings

High-precision three-dimensional (3D) positioning in dense urban non-line-of-sight (NLOS) environments benefits significantly from cooperation among multiple distributed base stations (BSs). However, forwarding raw CSI from multiple BSs to a central unit (CU) incurs prohibitive fronthaul overhead, which limits scalable cooperative positioning in practice. This paper proposes a learning-based edge-cloud cooperative positioning framework under limited-capacity fronthaul constraints. In the proposed architecture, a neural network is deployed at each BS to compress the locally estimated CSI into a quantized representation subject to a fixed fronthaul payload. The quantized CSI is transmitted to the CU, which performs cooperative 3D positioning by jointly processing the compressed CSI received from multiple BSs. The proposed framework adopts a two-stage training strategy consisting of self-supervised local training at the BSs and end-to-end joint training for positioning at the CU. Simulation results based on a 3.5~GHz 5G NR compliant urban ray-tracing scenario with six BSs and 20~MHz bandwidth show that the proposed method achieves a mean 3D positioning error of 0.48~m and a 90th-percentile error of 0.83~m, while reducing the fronthaul payload to 6.25% of lossless CSI forwarding. The achieved performance is close to that of cooperative positioning with full CSI exchange.

eess.SP

CMANet: Channel-Masked Attention Network for Cooperative Multi-Base-Station 3D Positioning

Achieving ubiquitous high-accuracy localization is crucial for next-generation wireless systems, yet remains challenging in multipath-rich urban environments. By exploiting the fine-grained multipath characteristics embedded in channel state information (CSI), more reliable and precise localization can be achieved. To address this, we present CMANet, a multi-BS cooperative positioning architecture that performs feature-level fusion of raw CSI using the proposed Channel Masked Attention (CMA) mechanism. The CMA encoder injects a physically grounded prior--per-BS channel gain--into the attention weights, thus emphasizing reliable links and suppressing spurious multipath. A lightweight LSTM decoder then treats subcarriers as a sequence to accumulate frequency-domain evidence into a final 3D position estimate. In a typical 5G NR-compliant urban simulation, CMANet achieves less than 0.5m median error and 1.0m 90th-percentile error, outperforming state-of-the-art benchmarks. Ablations verify the necessity of CMA and frequency accumulation. CMANet is edge-deployable and exemplifies an Integrated Sensing and Communication (ISAC)-aligned, cooperative paradigm for multi-BS CSI positioning.

eess.SP

Neural optimization of the most probable paths of 3D active Brownian particles

We develop a variational neural-network framework to determine the most probable path (MPP) of a 3D active Brownian particle (ABP) by directly minimizing the Onsager-Machlup integral (OMI). To obtain the OMI, we use the Onsager-Machlup variational principle for active systems and construct the Rayleighian of the ABP by including its active power. This approach reveals geometric transitions of the MPP from in-plane I- and U-shaped paths to 3D helical paths as the final time and net displacement are varied. We also demonstrate that the initial and final boundary conditions have a significant impact on the MPPs. Our results show that neural optimization combined with the Onsager-Machlup variational principle provides an efficient and versatile framework for exploring optimal transition pathways in active and nonequilibrium systems.

cond-mat.soft

Diffusive dynamics of charge regulated macro-ion solutions

Onsager's variational principle is generalized to address the diffusive dynamics of an electrolyte solution composed of charge-regulated macro-ions and counterions. The free energy entering the Rayleighian corresponds to the Poisson-Boltzmann theory augmented by the charge-regulation mechanism. The dynamical equations obtained by minimizing the Rayleighian include the classical Poisson-Nernst-Planck equations, the Debye-Falkenhagen equation, and their modifications in the presence of charge regulation. By analyzing the steady state, we show that the charge regulation has an important impact on the non-equilibrium macro-ion spatial distribution and their effective charge, deviating significantly from their equilibrium values. Our model, based on Onsager's variational principle offers a unified approach to the diffusive dynamics of electrolytes containing components that undergo various charge association/dissociation processes.

cond-mat.soft

Mammo-CLIP: Leveraging Contrastive Language-Image Pre-training (CLIP) for Enhanced Breast Cancer Diagnosis with Multi-view Mammography

Although fusion of information from multiple views of mammograms plays an important role to increase accuracy of breast cancer detection, developing multi-view mammograms-based computer-aided diagnosis (CAD) schemes still faces challenges and no such CAD schemes have been used in clinical practice. To overcome the challenges, we investigate a new approach based on Contrastive Language-Image Pre-training (CLIP), which has sparked interest across various medical imaging tasks. By solving the challenges in (1) effectively adapting the single-view CLIP for multi-view feature fusion and (2) efficiently fine-tuning this parameter-dense model with limited samples and computational resources, we introduce Mammo-CLIP, the first multi-modal framework to process multi-view mammograms and corresponding simple texts. Mammo-CLIP uses an early feature fusion strategy to learn multi-view relationships in four mammograms acquired from the CC and MLO views of the left and right breasts. To enhance learning efficiency, plug-and-play adapters are added into CLIP image and text encoders for fine-tuning parameters and limiting updates to about 1% of the parameters. For framework evaluation, we assembled two datasets retrospectively. The first dataset, comprising 470 malignant and 479 benign cases, was used for few-shot fine-tuning and internal evaluation of the proposed Mammo-CLIP via 5-fold cross-validation. The second dataset, including 60 malignant and 294 benign cases, was used to test generalizability of Mammo-CLIP. Study results show that Mammo-CLIP outperforms the state-of-art cross-view transformer in AUC (0.841 vs. 0.817, 0.837 vs. 0.807) on both datasets. It also surpasses previous two CLIP-based methods by 20.3% and 14.3%. This study highlights the potential of applying the finetuned vision-language models for developing next-generation, image-text-based CAD schemes of breast cancer.

cs.CV

Polariton microfluidics for nonreciprocal dragging and reconfigurable shaping of polaritons

Dielectric environment engineering is an efficient and general approach to manipulating polaritons. Liquids serving as surrounding media of polaritons have been used to shift polariton dispersions and tailor polariton wavefronts. However, those liquid-based methods have so far been limited to their static states, not fully unleashing the promises offered by the mobility of liquids. Here, we propose a microfluidic strategy for polariton manipulation by merging polaritonics with microfluidics. The diffusion of fluids causes gradient refractive indices over microchannels, which breaks the symmetry of polariton dispersions and realizes the non-reciprocal dragging of polaritons. Based on polariton microfluidics, we also design a set of on-chip polaritonic elements to actively shape polaritons, including planar lenses, off-axis lenses, Janus lenses, bends, and splitters. Our strategy expands the toolkit for the manipulation of polaritons at the subwavelength scale and possesses potential in the fields of polariton biochemistry and molecular sensing.

physics.optics

Charge Regulation of Polyelectrolyte Gels: Swelling Transition

We study the effects of charge-regulated acid/base equilibrium on the swelling of polyelectrolyte gels, by considering a combination of the Poisson-Boltzmann theory and a two-site charge-regulation model based on the Langmuir adsorption isotherm. By exploring the volume change as a function of salt concentration for both nano-gels and micro-gels, we identify conditions where the gel volume exhibits a discontinuous swelling transition. This transition is driven exclusively by the charge-regulation mechanism and is characterized by a closed-loop phase diagram. Our predictions can be tested experimentally for polypeptide gels.

cond-mat.soft

Universality in the dynamics of vesicle translocation through a hole

We analyze the translocation process of a spherical vesicle, made of membrane and incompressible fluid, through a hole smaller than the vesicle size, driven by pressure difference $ΔP$. We show that such a vesicle shows certain universal characteristics which is independent of the details of the membrane elasticity; (i) there is a critical pressure $ΔP_{\rm c}$ below which no translocation occurs, (ii) $ΔP_{\rm c}$ decreases to zero as the vesicle radius $R_0$ approaches the hole radius $a$, satisfying the scaling relation $ΔP_{\rm c} \sim (R_0 - a)^{3/2}$, and (iii) the translocation time $τ$ diverges as $ΔP$ decreases to $ΔP_{\rm c}$, satisfying the scaling relation $τ\sim (ΔP -ΔP_{\rm c})^{-1/2}$.

cond-mat.soft

Transformers Improve Breast Cancer Diagnosis from Unregistered Multi-View Mammograms

Deep convolutional neural networks (CNNs) have been widely used in various medical imaging tasks. However, due to the intrinsic locality of convolution operation, CNNs generally cannot model long-range dependencies well, which are important for accurately identifying or mapping corresponding breast lesion features computed from unregistered multiple mammograms. This motivates us to leverage the architecture of Multi-view Vision Transformers to capture long-range relationships of multiple mammograms from the same patient in one examination. For this purpose, we employ local Transformer blocks to separately learn patch relationships within four mammograms acquired from two-view (CC/MLO) of two-side (right/left) breasts. The outputs from different views and sides are concatenated and fed into global Transformer blocks, to jointly learn patch relationships between four images representing two different views of the left and right breasts. To evaluate the proposed model, we retrospectively assembled a dataset involving 949 sets of mammograms, which include 470 malignant cases and 479 normal or benign cases. We trained and evaluated the model using a five-fold cross-validation method. Without any arduous preprocessing steps (e.g., optimal window cropping, chest wall or pectoral muscle removal, two-view image registration, etc.), our four-image (two-view-two-side) Transformer-based model achieves case classification performance with an area under ROC curve (AUC = 0.818), which significantly outperforms AUC = 0.784 achieved by the state-of-the-art multi-view CNNs (p = 0.009). It also outperforms two one-view-two-side models that achieve AUC of 0.724 (CC view) and 0.769 (MLO view), respectively. The study demonstrates the potential of using Transformers to develop high-performing computer-aided diagnosis schemes that combine four mammograms.

cs.CV

Enhanced electro-actuation in dielectric elastomers: the non-linear effect of free ions

Plasticized poly(vinyl chloride) (PVC) is a jelly-like soft dielectric material that attracted substantial interest recently as a new type of electro-active polymers. Under electric fields of several hundred Volt/mm, PVC gels undergo large deformations. These gels can be used as artificial muscles and other soft robotic devices, with striking deformation behavior that is quite different from conventional dielectric elastomers. Here, we present a simple model for the electro-activity of PVC gels, and show a non-linear effect of free ions on its dielectric behaviors. It is found that their particular deformation behavior is due to an electro-wetting effect and to a change in their interfacial tension. In addition, we derive analytical expressions for the surface tension as well as for the apparent dielectric constant of the gel. The theory indicates that the size of the mobile free ions has a crucial role in determining the electro-induced deformation, opening up the way to novel and innovative designs of electro-active gel actuators.

cond-mat.soft

Recent advances and clinical applications of deep learning in medical image analysis

Deep learning has received extensive research interest in developing new medical image processing algorithms, and deep learning based models have been remarkably successful in a variety of medical imaging tasks to support disease detection and diagnosis. Despite the success, the further improvement of deep learning models in medical image analysis is majorly bottlenecked by the lack of large-sized and well-annotated datasets. In the past five years, many studies have focused on addressing this challenge. In this paper, we reviewed and summarized these recent studies to provide a comprehensive overview of applying deep learning methods in various medical image analysis tasks. Especially, we emphasize the latest progress and contributions of state-of-the-art unsupervised and semi-supervised deep learning in medical image analysis, which are summarized based on different application scenarios, including classification, segmentation, detection, and image registration. We also discuss the major technical challenges and suggest the possible solutions in future research efforts.

cs.CV

Virtual Adversarial Training for Semi-supervised Breast Mass Classification

This study aims to develop a novel computer-aided diagnosis (CAD) scheme for mammographic breast mass classification using semi-supervised learning. Although supervised deep learning has achieved huge success across various medical image analysis tasks, its success relies on large amounts of high-quality annotations, which can be challenging to acquire in practice. To overcome this limitation, we propose employing a semi-supervised method, i.e., virtual adversarial training (VAT), to leverage and learn useful information underlying in unlabeled data for better classification of breast masses. Accordingly, our VAT-based models have two types of losses, namely supervised and virtual adversarial losses. The former loss acts as in supervised classification, while the latter loss aims at enhancing model robustness against virtual adversarial perturbation, thus improving model generalizability. To evaluate the performance of our VAT-based CAD scheme, we retrospectively assembled a total of 1024 breast mass images, with equal number of benign and malignant masses. A large CNN and a small CNN were used in this investigation, and both were trained with and without the adversarial loss. When the labeled ratios were 40% and 80%, VAT-based CNNs delivered the highest classification accuracy of 0.740 and 0.760, respectively. The experimental results suggest that the VAT-based CAD scheme can effectively utilize meaningful knowledge from unlabeled data to better classify mammographic breast mass images.

cs.CV