SearcharxivSearch

arXiv subjects

Bin Tang

Publications and source records attributed to Bin Tang.

At least 19 recordsLinked to original sources

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI safety. We present Yuvion VL, a family of multimodal large language models purpose-built for content and AI safety, with both instruction-tuned and reasoning-oriented variants. Yuvion VL addresses this gap by treating safety as an inherently adversarial and multimodal problem and designing the entire pipeline around adversarial robustness. For data construction, we develop an automated pipeline integrating adversarial-aware data synthesis with multi-stage quality control, producing large-scale, high-quality multimodal samples augmented with domain knowledge and reasoning annotations. For training, we adopt a three-stage pipeline that includes continued pretraining for risk-concept cross-modal alignment, instruct post-training for production-grade safety tasks, and reasoning post-training for enhanced interpretability and performance in complex tasks. We further introduce Confuse-then-Contrast Fine-Tuning, a contrastive framework that mines model-specific confusions and constructs multi-image contrastive groups to enforce explicit discrimination of fine-grained visual-semantic elements, enabling the model to distinguish between visually similar cases with different safety implications in adversarial safety tasks. To support rigorous evaluation, we further introduce Yuvion VL RiskEval (YVRE), a collection of benchmarks covering diverse open and internal evaluations, with a focus on content and AI safety, adversarial robustness, and real-world capability requirements. Experiments show that Yuvion VL-32B achieves industry-leading safety performance, surpassing comparably sized open-source models and best closed-source commercial models, while maintaining comparable general capabilities.

cs.CV

Large deviation principles for the stationary solutions and invariant measures of a class of SPDE with locally monotone coefficients

We establish the well-posedness of stationary solutions for a class of SPDEs with locally monotone coefficients, and prove the Freidlin--Wentzell large deviation principle (LDP) for these stationary solutions. The LDP for the associated invariant measures then follows via the contraction principle, avoiding the need to construct the quasi-potential and verify the Dembo--Zeitouni uniform LDP over bounded sets. By working directly with stationary solutions, we bypass these technical difficulties, thereby providing a more general and flexible framework that is adapted to additive noise, multiplicative noise, and transport-type noise. As applications, our results cover a range of SPDEs, including the stochastic reaction-diffusion equations, stochastic 1D viscous Burgers equation, stochastic 2D Navier--Stokes equations, stochastic 2D magneto-hydrodynamic equations and stochastic 3D hyper-dissipative Navier--Stokes equations.

math.PR

Joint Time-Phase Synchronization for Distributed Sensing Networks via Feature-Level Hyper-Plane Regression

Achieving coherent integration in distributed Internet of Things (IoT) sensing networks requires precise synchronization to jointly compensate clock offsets and radio-frequency (RF) phase errors. Conventional two-step protocols suffer from time-phase coupling, where residual timing offsets degrade phase coherence. This paper proposes a generalized hyper-plane regression (GHR) framework for joint calibration by transforming coupled spatiotemporal phase evolution into a unified regression model, enabling effective parameter decoupling. To support resource-constrained IoT edge nodes, a feature-level distributed architecture is developed. By adopting a linear frequency-modulated (LFM) waveform, the model order is reduced, yielding linear computational complexity. In addition, a unidirectional feature transmission mechanism eliminates the communication overhead of bidirectional timestamp exchange, making the approach suitable for resource-constrained IoT networks. Simulation results demonstrate reliable picosecond-level synchronization accuracy under severe noise across kilometer-scale distributed IoT sensing networks.

eess.SP

Geometric Direction Finding on Dynamic Manifolds: Unambiguous DOA Estimation for Spatially Undersampled UWB Arrays

Traditional Direction of Arrival (DOA) estimation methods struggle to simultaneously address three physical constraints in Ultra-Wideband (UWB) electromagnetic sensing: spatial undersampling, asynchronous array phase, and beam squint. Existing solutions treat these issues in isolation, leading to limited performance in complex scenarios. This paper proposes a novel dynamic manifold perspective, which models UWB signal observations as a continuous manifold curve in a high-dimensional space driven by temporal evolution and array topology. We theoretically demonstrate that the DOA can be uniquely determined solely by the geometric shape of the manifold, rather than the absolute arrival phase. Based on this perspective, we construct a geometric parameter system comprising extrinsic and intrinsic parameters, along with a corresponding DOA estimation framework. Extrinsic vector parameters serve as a dynamic extension of traditional array processing, effectively expanding the degrees of freedom to suppress grating lobes. Intrinsic scalar invariants provide a new geometric perspective independent of traditional phase models, offering intrinsic robustness against array channel phase errors. Simulation results show that the derived analytical expressions for geometric parameters are highly consistent with numerical truths. The proposed framework not only completely eliminates spatial ambiguity in sparse arrays but also achieves high-precision direction finding under conditions with calibration-free phase errors.

eess.SP

Enhance and Reuse: A Dual-Mechanism Approach to Boost Deep Forest for Label Distribution Learning

Label distribution learning (LDL) requires the learner to predict the degree of correlation between each sample and each label. To achieve this, a crucial task during learning is to leverage the correlation among labels. Deep Forest (DF) is a deep learning framework based on tree ensembles, whose training phase does not rely on backpropagation. DF performs in-model feature transform using the prediction of each layer and achieves competitive performance on many tasks. However, its exploration in the field of LDL is still in its infancy. The few existing methods that apply DF to the field of LDL do not have effective ways to utilize the correlation among labels. Therefore, we propose a method named Enhanced and Reused Feature Deep Forest (ERDF). It mainly contains two mechanisms: feature enhancement exploiting label correlation and measure-aware feature reuse. The first one is to utilize the correlation among labels to enhance the original features, enabling the samples to acquire more comprehensive information for the task of LDL. The second one performs a reuse operation on the features of samples that perform worse than the previous layer on the validation set, in order to ensure the stability of the training process. This kind of Enhance-Reuse pattern not only enables samples to enrich their features but also validates the effectiveness of their new features and conducts a reuse process to prevent the noise from spreading further. Experiments show that our method outperforms other comparison algorithms on six evaluation metrics.

cs.LG

When Vision-Language Model (VLM) Meets Beam Prediction: A Multimodal Contrastive Learning Framework

As the real propagation environment becomes in creasingly complex and dynamic, millimeter wave beam prediction faces huge challenges. However, the powerful cross modal representation capability of vision-language model (VLM) provides a promising approach. The traditional methods that rely on real-time channel state information (CSI) are computationally expensive and often fail to maintain accuracy in such environments. In this paper, we present a VLM-driven contrastive learning based multimodal beam prediction framework that integrates multimodal data via modality-specific encoders. To enforce cross-modal consistency, we adopt a contrastive pretraining strategy to align image and LiDAR features in the latent space. We use location information as text prompts and connect it to the text encoder to introduce language modality, which further improves cross-modal consistency. Experiments on the DeepSense-6G dataset show that our VLM backbone provides additional semantic grounding. Compared with existing methods, the overall distance-based accuracy score (DBA-Score) of 0.9016, corresponding to 1.46% average improvement.

eess.SP

A broadband platform to search for hidden photons

The optical behavior of a structure consisting of graphene sheets embedded in media was studied, and the differences between the structure and ordinary birefringent crystal, double zero-reflectance point, were identified. We showed the changes in the optical behavior of the structure due to the existence of hidden photons. When a radiation illuminates the structure, only $\omega^2/\omega_p^2>1+\frac{m_X^2 c^4 \chi^2}{\epsilon_r\hbar^2\omega_p^2}$ can propagate through the structure. This provides a broadband platform for detecting hidden photons, where the sensitivity increases with the mass of the hidden photon.In contrast, if the mass of hidden photon is small, one can use a method similar to the light-shining-through-thin-wall technique. The structure is a platform to actively search for hidden photons since the operating point of the structure does not have to match the mass shell of hidden photons.

physics.optics

Pose as a Modality: A Psychology-Inspired Network for Personality Recognition with a New Multimodal Dataset

In recent years, predicting Big Five personality traits from multimodal data has received significant attention in artificial intelligence (AI). However, existing computational models often fail to achieve satisfactory performance. Psychological research has shown a strong correlation between pose and personality traits, yet previous research has largely ignored pose data in computational models. To address this gap, we develop a novel multimodal dataset that incorporates full-body pose data. The dataset includes video recordings of 287 participants completing a virtual interview with 36 questions, along with self-reported Big Five personality scores as labels. To effectively utilize this multimodal data, we introduce the Psychology-Inspired Network (PINet), which consists of three key modules: Multimodal Feature Awareness (MFA), Multimodal Feature Interaction (MFI), and Psychology-Informed Modality Correlation Loss (PIMC Loss). The MFA module leverages the Vision Mamba Block to capture comprehensive visual features related to personality, while the MFI module efficiently fuses the multimodal features. The PIMC Loss, grounded in psychological theory, guides the model to emphasize different modalities for different personality dimensions. Experimental results show that the PINet outperforms several state-of-the-art baseline models. Furthermore, the three modules of PINet contribute almost equally to the model's overall performance. Incorporating pose data significantly enhances the model's performance, with the pose modality ranking mid-level in importance among the five modalities. These findings address the existing gap in personality-related datasets that lack full-body pose data and provide a new approach for improving the accuracy of personality prediction models, highlighting the importance of integrating psychological insights into AI frameworks.

cs.CV

Enhance Learning Efficiency of Oblique Decision Tree via Feature Concatenation

Oblique Decision Tree (ODT) separates the feature space by linear projections, as opposed to the conventional Decision Tree (DT) that forces axis-parallel splits. ODT has been proven to have a stronger representation ability than DT, as it provides a way to create shallower tree structures while still approximating complex decision boundaries. However, its learning efficiency is still insufficient, since the linear projections cannot be transmitted to the child nodes, resulting in a waste of model parameters. In this work, we propose an enhanced ODT method with Feature Concatenation (\texttt{FC-ODT}), which enables in-model feature transformation to transmit the projections along the decision paths. Theoretically, we prove that our method enjoys a faster consistency rate w.r.t. the tree depth, indicating that our method possesses a significant advantage in generalization performance, especially for shallow trees. Experiments show that \texttt{FC-ODT} can outperform the other state-of-the-art decision trees with a limited tree depth.

cs.LG

Improving Multi-Label Contrastive Learning by Leveraging Label Distribution

In multi-label learning, leveraging contrastive learning to learn better representations faces a key challenge: selecting positive and negative samples and effectively utilizing label information. Previous studies selected positive and negative samples based on the overlap between labels and used them for label-wise loss balancing. However, these methods suffer from a complex selection process and fail to account for the varying importance of different labels. To address these problems, we propose a novel method that improves multi-label contrastive learning through label distribution. Specifically, when selecting positive and negative samples, we only need to consider whether there is an intersection between labels. To model the relationships between labels, we introduce two methods to recover label distributions from logical labels, based on Radial Basis Function (RBF) and contrastive loss, respectively. We evaluate our method on nine widely used multi-label datasets, including image and vector datasets. The results demonstrate that our method outperforms state-of-the-art methods in six evaluation metrics.

cs.LG

Large deviation principle for the stationary solutions of stochastic functional differential equations with infinite delay

We investigate the large deviation principle (LDP) of the stationary solutions of stochastic functional differential equations (SFDEs) with infinite delay under small random perturbation. First, we demonstrate the existence and uniqueness of the corresponding stationary solutions. Second, by the weak convergence approach, we show the uniform large deviation principle for the solution maps, and then prove the LDP for stationary solutions. Furthermore, we obtain the LDP for invariant measures of SFDEs through the LDP for stationary solutions and the contraction principle.

math.PR

SSFam: Scribble Supervised Salient Object Detection Family

Scribble supervised salient object detection (SSSOD) constructs segmentation ability of attractive objects from surroundings under the supervision of sparse scribble labels. For the better segmentation, depth and thermal infrared modalities serve as the supplement to RGB images in the complex scenes. Existing methods specifically design various feature extraction and multi-modal fusion strategies for RGB, RGB-Depth, RGB-Thermal, and Visual-Depth-Thermal image input respectively, leading to similar model flood. As the recently proposed Segment Anything Model (SAM) possesses extraordinary segmentation and prompt interactive capability, we propose an SSSOD family based on SAM, named SSFam, for the combination input with different modalities. Firstly, different modal-aware modulators are designed to attain modal-specific knowledge which cooperates with modal-agnostic information extracted from the frozen SAM encoder for the better feature ensemble. Secondly, a siamese decoder is tailored to bridge the gap between the training with scribble prompt and the testing with no prompt for the stronger decoding ability. Our model demonstrates the remarkable performance among combinations of different modalities and refreshes the highest level of scribble supervised methods and comes close to the ones of fully supervised methods. https://github.com/liuzywen/SSFam

cs.CV

Constrained motion of self-propelling eccentric disks linked by a spring

It has been supposed that the interplay of elasticity and activity plays a key role in triggering the non-equilibrium behaviors in biological systems. However, the experimental model system is missing to investigate the spatiotemporally dynamical phenomena. Here, a model system of an active chain, where active eccentric-disks are linked by a spring, is designed to study the interplay of activity, elasticity, and friction. Individual active chain exhibits longitudinal and transverse motion, however, it starts to self-rotate when pinning one end, and self-beats when clamping one end. Additionally, our eccentric-disk model can qualitatively reproduce such behaviors and explain the unusual self-rotation of the first disk around its geometric center. Further, the structure and dynamics of long chains were studied via simulations without steric interactions. It was found that hairpin conformation emerges in free motion, while in the constrained motions, the rotational and beating frequencies scale with the flexure number (the ratio of self-propelling force to bending rigidity), ~4/3. Scaling analysis suggests that it results from the balance between activity and energy dissipation. Our findings show that topological constraints play a vital role in non-equilibrium synergy behavior.

cond-mat.soft

Escape of an Active Ring from an Attractive Surface: Behaving Like a Self-Propelled Brownian Particle

Escape of active agents from metastable states is of great interest in statistical and biological physics. In this study, we investigate the escape of a flexible active ring, composed of active Brownian particles, from a flat attractive surface using Brownian dynamics simulations. To systematically explore the effects of activity, persistence time, and the shape of attractive potentials, we calculate escape time and effective temperature. We observe two distinct escape mechanisms: Kramers-like thermal activation at small persistence times and the maximal force problem at large persistence time, where escape time is determined by persistence time. The escape time explicitly depends on the shape of the potential barrier at high activity and large persistence time. Moreover, when the propulsion force is biased along the ring's contour, escape becomes more difficult and is primarily driven by thermal noise. Our findings highlight that, despite its intricate configuration, the active ring can be effectively modeled as a self-propelled Brownian particle when studying its escape from a smooth surface.

cond-mat.soft

An elementary approach to mixing and dissipation enhancement by transport noise

We investigate the mixing properties of solutions to the stochastic transport equation $d u= \circ d W \cdot\nabla u$, where the driving noise $W(t,x)$ is white in time, colored and divergence-free in space. Furthermore, we prove the dissipation enhancement in the presence of a small viscous term. Applying our results, we also derive the mixing properties for a regularized stochastic 2D Euler equation.

math.PR

A Fast Power Spectrum Sensing Solution for Generalized Coprime Sampling

The growing scarcity of spectrum resources, wideband spectrum sensing is required to process a prohibitive volume of data at a high sampling rate. For some applications, spectrum estimation only requires second-order statistics. In this case, a fast power spectrum sensing solution is proposed based on the generalized coprime sampling. By exploring the sensing vector inherent structure, the autocorrelation sequence of inputs can be reconstructed from sub-Nyquist samples by only utilizing the parallel Fourier transform and simple multiplication operations. Thus, it takes less time than the state-of-the-art methods while maintaining the same performance, and it achieves higher performance than the existing methods within the same execution time, without the need for pre-estimating the number of inputs. Furthermore, the influence of the model mismatch has only a minor impact on the estimation performance, which allows for more efficient use of the spectrum resource in a distributed swarm scenario. Simulation results demonstrate the low complexity in sampling and computation, making it a more practical solution for real-time and distributed wideband spectrum sensing applications.

eess.SP

Multiscale Motion-Aware and Spatial-Temporal-Channel Contextual Coding Network for Learned Video Compression

Recently, learned video compression has achieved exciting performance. Following the traditional hybrid prediction coding framework, most learned methods generally adopt the motion estimation motion compensation (MEMC) method to remove inter-frame redundancy. However, inaccurate motion vector (MV) usually lead to the distortion of reconstructed frame. In addition, most approaches ignore the spatial and channel redundancy. To solve above problems, we propose a motion-aware and spatial-temporal-channel contextual coding based video compression network (MASTC-VC), which learns the latent representation and uses variational autoencoders (VAEs) to capture the characteristics of intra-frame pixels and inter-frame motion. Specifically, we design a multiscale motion-aware module (MS-MAM) to estimate spatial-temporal-channel consistent motion vector by utilizing the multiscale motion prediction information in a coarse-to-fine way. On the top of it, we further propose a spatial-temporal-channel contextual module (STCCM), which explores the correlation of latent representation to reduce the bit consumption from spatial, temporal and channel aspects respectively. Comprehensive experiments show that our proposed MASTC-VC is surprior to previous state-of-the-art (SOTA) methods on three public benchmark datasets. More specifically, our method brings average 10.15\% BD-rate savings against H.265/HEVC (HM-16.20) in PSNR metric and average 23.93\% BD-rate savings against H.266/VVC (VTM-13.2) in MS-SSIM metric.

eess.IV

Wideband Power Spectrum Sensing: a Fast Practical Solution for Nyquist Folding Receiver

The limited availability of spectrum resources has been growing into a critical problem in wireless communications, remote sensing, and electronic surveillance, etc. To address the high-speed sampling bottleneck of wideband spectrum sensing, a fast and practical solution of power spectrum estimation for Nyquist folding receiver (NYFR) is proposed in this paper. The NYFR architectures is can theoretically achieve the full-band signal sensing with a hundred percent of probability of intercept. But the existing algorithm is difficult to realize in real-time due to its high complexity and complicated calculations. By exploring the sub-sampling principle inherent in NYFR, a computationally efficient method is introduced with compressive covariance sensing. That can be efficient implemented via only the non-uniform fast Fourier transform, fast Fourier transform, and some simple multiplication operations. Meanwhile, the state-of-the-art power spectrum reconstruction model for NYFR of time-domain and frequency-domain is constructed in this paper as a comparison. Furthermore, the computational complexity of the proposed method scales linearly with the Nyquist-rate sampled number of samples and the sparsity of spectrum occupancy. Simulation results and discussion demonstrate that the low complexity in sampling and computation is a more practical solution to meet the real-time wideband spectrum sensing applications.

eess.SP