SearcharxivSearch

arXiv subjects

Yining Sun

Publications and source records attributed to Yining Sun.

At least 19 recordsLinked to original sources

Classification of Novikov-Poisson Algebras and Their Applications

In this paper, we give a complete classification, up to isomorphism, of 3-dimensional complex Novikov--Poisson algebras. As an application of the classification, we further prove that every 3-dimensional complex transposed Poisson algebra can be obtained from a Novikov--Poisson algebra except the Lie algebra $\mathfrak{sl}_2(\mathbb C)$. Consequently, Sartayev's conjecture holds in dimension 3, that is, every 3-dimensional complex transposed Poisson algebra is special when regarded as a GD algebra.

math.RA

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks. However, we identify two primary limitations in existing works: 1) The multi-layered characteristics of harmful videos are overlooked. Existing benchmarks predominantly formulate evaluation as a binary classification task, failing to capture implicit or deep contextual harms. 2) Explanatory rationales are completely absent. Current frameworks measure exclusively whether a model flags a video correctly rather than explaining why, turning evaluation into a black box where models can succeed through superficial shortcuts. To address these problems, we present HarmVideoBench, a multi-layered diagnostic benchmark comprising 1,379 videos paired with 4,137 multiple-choice questions. HarmVideoBench benchmarks three hierarchical dimensions: Observable Evidence, Clip-Internal Meaning, and Beyond-Clip Reasoning, aiming to evaluate models' deep understanding beyond surface cues with carefully balanced and curated samples. We evaluate 19 leading models on HarmVideoBench to assess their multidimensional understanding of harmful videos. Moreover, we introduce BCR, a benchmark-aligned method that predicts reasoning boundaries and dynamically retrieves context only when needed. Experimental results show that BCR substantially improves the base model's performance in harmful video understanding, raising the macro average from 61.7 percent to a state-of-the-art 84.4 percent.

cs.CV

VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as arrows, sketches, and emojis orchestrate complex video dynamics with unprecedented controllability. However, these seemingly innocuous static cues can be interpreted by models as executable temporal instructions, unfolding into harmful actions in the generated videos. Despite the severity of this threat, existing safety benchmarks remain predominantly focused on text-based and content-only image-based jailbreaks, leaving implicit visual prompt attacks insufficiently explored. To bridge this gap, we present VVA-Bench, the first systematic benchmark for evaluating video generation safety under categorized vision-centric prompt attacks. Extensive experiments on VVA-Bench demonstrate that state-of-the-art models are highly susceptible to such attacks, with Attack Success Rates (ASR) reaching 100.0\% on Wan 2.7 and 74.8\% on Veo 3.1. To mitigate these risks, we propose VPA-Guard, a retrieval-augmented and self-evolving defense framework. By leveraging few-shot reasoning to identify latent malicious intents, our method reduces the attack ASR by 44.2\% and the harmfulness score by 73.4\% on average, while maintaining the model's utility for legitimate user edits. Our work provides both a rigorous benchmark and an effective defense strategy to advance safe and socially responsible multimodal generation.

cs.CV

Rapid intermediate-mass black hole formation via runaway mergers of black holes

Observations indicate that supermassive black holes (SMBHs) in high-redshift galaxies formed on timescales far shorter than classical growth models allow. One hypothesis suggests intermediate-mass black hole (IMBH) seeds as an efficient growth channel. Using N-body simulations, we demonstrate that in dense stellar-mass black hole (BH) clusters ($\ge 5\times10^9 M_{\odot}/{\rm pc}^3$), runaway gravitational-wave binary BH (BBH) mergers can produce a $\sim 10^3 M_\odot$ IMBH within 10 Myr from the formation of the BH subsystem. This scenario is simple and avoids large uncertainties regarding stellar mergers and evolution in the IMBH formation via very massive stars channel. We find that the runaway GW-merger mechanism relies on hard BBH formation through a chain of exchanged soft BBHs with accumulated hardening, which is far more efficient than three-body scattering. We analyze how IMBH formation depends on cluster density, total mass, initial mass function, and stellar halo potential. We find that due to cluster expansion, the systems forming IMBHs have densities consistent with present-day nuclear star clusters, such as those in the Milky Way and M33. Furthermore, we show that IMBH spin remains low due to repeated mergers, and we estimate the rate of GW190521 and GW231123-like events within the first 100 Myr to be $2.27-247.52$ and $3.23-63.63 $ per Gyr per cluster.

astro-ph.GA

Gelfand--Dorfman Algebras: Nilpotency, Solvability, Construction and Classification

In this paper, we characterize the nilpotency and solvability of Gelfand--Dorfman (GD) algebras. In contrast with Poisson algebras and transposed Poisson algebras, we give examples show the nilpotency and solvability of a GD algebra are not determined by the nilpotency and solvability of its underlying algebras. To obtain more examples of special and non-special GD algebras, we give several construction methods and determine whether the resulting algebras are special. Futhermore, we study GD algebra structures on simple Lie algebras. We provide examples demonstrating that GD algebra structures on simple Lie algebras are not necessarily trivial, distinguishing them from Poisson and transposed Poisson algebras. Finally, we provide a complete algebraic classification of low-dimensional complex GD algebras, and determine their nilpotency, solvability and speciality.

math.RA

$A$-Generalized Hessian pre-Lie algebras and $A$-Generalized Yang--Baxter Equations

Inspired by the problem of constructing ($\omega$-)pre-Lie algebra structures on the dual space of a pre-Lie algebra, we introduce the \(A\)-generalized Yang--Baxter equation as a generalization of the Yang--Baxter equation of pre-Lie algebras. We study its symmetric solutions through \(A\)-generalized Hessian pre-Lie algebras and split these solutions into two types. We further consider factorizable solutions of this equation and establish a one-to-one correspondence between them and generalized quadratic Rota--Baxter pre-Lie algebras of nonzero weight. By studying the structure of these algebras, we find all factorizable solutions. Finally, we study the structure of \(A\)-generalized Hessian pre-Lie algebras. In particular, we obtain a structural description via central and double extensions and classify low-dimensional non-trivial \(A\)-generalized Hessian pre-Lie algebras.

math-ph

The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMs

Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) frequently exhibit a critical failure mode: they achieve high performance on in-distribution benchmarks while demonstrating brittle reasoning capabilities on out-of-distribution (OOD) tasks. We term this phenomenon Reward-Induced Manifold Collapse. We establish a theoretical framework bridging Structural Causal Models (SCM) and the Information Bottleneck (IB) principle to explain this paradox. We define reasoning as a high-complexity causal process and shortcut learning as the exploitation of low-complexity spurious correlations. Under the implicit inductive bias of Stochastic Gradient Descent (SGD), models optimized for outcome rewards are biased toward shortcut solutions whenever the training distribution allows for a ``Markovian Screening'' of the true causal mechanism. We derive a new generalization bound based on Semantic Coverage Measure ($\eta$) rather than sample size, showing why data scaling on homogeneous distributions may fail to correct reasoning flaws. We also show that Process Reward Models (PRMs) function as Topological Filters, enforcing step-wise mutual information constraints that render the low-complexity shortcut manifold inadmissible. These results provide a mathematical grounding for the role of process supervision beyond simple credit assignment.

cs.LG

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual inputs such as marks, arrows, and visual-text prompts. While this paradigm greatly expands usability, it also introduces a critical and underexplored safety risk: the attack surface itself becomes visual. In this work, we propose Vision-Centric Jailbreak Attack (VJA), the first visual-to-visual jailbreak attack that conveys malicious instructions purely through visual inputs. To systematically study this emerging threat, we introduce IESBench, a safety-oriented benchmark for image editing models. Extensive experiments on IESBench demonstrate that VJA effectively compromises state-of-the-art commercial models, achieving attack success rates of up to 80.9% on Nano Banana Pro and 70.1% on GPT-Image-1.5. To mitigate this vulnerability, we propose a training-free defense based on introspective multimodal reasoning, which substantially improves the safety of poorly aligned models to a level comparable with commercial systems, without auxiliary guard models and with negligible computational overhead. Our findings expose new vulnerabilities, provide both a benchmark and practical defense to advance safe and trustworthy modern image editing systems. Warning: This paper contains offensive images created by large image editing models.

cs.CV

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs).Despite recent benchmark gains reported for RL, we find that training on these benchmarks' training sets achieves nearly the same performance as training directly on the test sets, suggesting that the benchmarks cannot reliably separate further progress.To study this phenomenon, we introduce a diagnostic suite and the Oracle Performance Gap (OPG) metric that quantifies the performance difference between training on the train split versus the test split of a benchmark. We further analyze this phenomenon with stress tests and find that, despite strong benchmark scores, existing RL methods struggle to generalize across distribution shifts, varying levels of difficulty, and counterfactual scenarios: shortcomings that current benchmarks fail to reveal.We conclude that current benchmarks are insufficient for evaluating generalization and propose three core principles for designing more faithful benchmarks: sufficient difficulty, balanced evaluation, and distributional robustness.

cs.LG

$\omega$-Lie bialgebras and $\omega$-Yang-Baxter equation

In this paper, we introduce the definition of multiplicative $\omega$-Lie bialgebra, which is equivalent to the Manin triples and matched pairs. We also study the $\omega$-Yang-Baxter equation and Yang-Baxter $\omega$-Lie bialgebra. The skew-symmetric solutions of the $\omega$-Yang-Baxter equation can be used to construct Yang-Baxter $\omega$-Lie bialgebra. We further introduce the concept of the $\omega$-$\mathcal{O}$-operator, which can be constructed from a left-symmetric algebras, and based on the $\omega$-$\mathcal{O}$-operator, we construct skew-symmetric solutions to the $\omega$-Yang--Baxter equation.

math.RA

CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models

Autoregressive large vision--language models (LVLMs) interface video and language by projecting video features into the LLM's embedding space as continuous visual token embeddings. However, it remains unclear where temporal evidence is represented and how it causally influences decoding. To address this gap, we present CircuitProbe, a circuit-level analysis framework that dissects the end-to-end video-language pathway through two stages: (i) Visual Auditing, which localizes object semantics within the projected video-token sequence and reveals their causal necessity via targeted ablations and controlled substitutions; and (ii) Semantic Tracing, which uses logit-lens probing to track the layer-wise emergence of object and temporal concepts, augmented with temporal frame interventions to assess sensitivity to temporal structure. Based on the resulting analysis, we design a targeted surgical intervention that strictly follows our observations: identifying temporally specialized attention heads and selectively amplifying them within the critical layer interval revealed by Semantic Tracing. This analysis-driven intervention yields consistent improvements (up to 2.4% absolute) on the temporal-heavy TempCompass benchmark, validating the correctness, effectiveness, and practical value of the proposed circuit-level analysis for temporal understanding in LVLMs.

cs.CV

FRAME: Feedback-Refined Agent Methodology for Enhancing Medical Research Insights

The automation of scientific research through large language models (LLMs) presents significant opportunities but faces critical challenges in knowledge synthesis and quality assurance. We introduce Feedback-Refined Agent Methodology (FRAME), a novel framework that enhances medical paper generation through iterative refinement and structured feedback. Our approach comprises three key innovations: (1) A structured dataset construction method that decomposes 4,287 medical papers into essential research components through iterative refinement; (2) A tripartite architecture integrating Generator, Evaluator, and Reflector agents that progressively improve content quality through metric-driven feedback; and (3) A comprehensive evaluation framework that combines statistical metrics with human-grounded benchmarks. Experimental results demonstrate FRAME's effectiveness, achieving significant improvements over conventional approaches across multiple models (9.91% average gain with DeepSeek V3, comparable improvements with GPT-4o Mini) and evaluation dimensions. Human evaluation confirms that FRAME-generated papers achieve quality comparable to human-authored works, with particular strength in synthesizing future research directions. The results demonstrated our work could efficiently assist medical research by building a robust foundation for automated medical research paper generation while maintaining rigorous academic standards.

cs.CL

Pair-Instability Gap Black Holes in Population III Star Clusters: Pathways, Dynamics, and Gravitational Wave Implications

The detection of the gravitational wave (GW) event GW190521 raises questions about the formation of black holes within the pair-instability mass gap (PIBHs). We propose that Population III (Pop III) star clusters significantly contribute to events similar to GW190521. We perform $N$-body simulations and find that PIBHs can form from stellar collisions or binary black hole (BBH) mergers, with the latter accounting for 90\% of the contributions. Due to GW recoil during BBH mergers, approximately 10-50% of PIBHs formed via BBH mergers escape from clusters, depending on black hole spins and cluster escape velocities. The remaining PIBHs can participate in secondary and multiple BBH formation events, contributing to GW events. Assuming Pop III stars form in massive clusters (initially 100,000 $M_\odot$) with a top-heavy initial mass function, the average merger rates for GW events involving PIBHs with 0% and 100% primordial binaries are $0.005$ and $0.017$ $\text{yr}^{-1} \text{Gpc}^{-3}$, respectively, with maximum values of $0.030$ and $0.106$ $\text{yr}^{-1} \text{Gpc}^{-3}$. If Pop III stars form in low-mass clusters (initial mass of $1000M_\odot$ and $10000 M_\odot$), the merger rate is comparable with a 100% primordial binary fraction but significantly lower without primordial binaries. We also calculate the characteristic strains of the GW events in our simulations and find that about 43.4% (LISA) 97.8% (Taiji) and 66.4% (Tianqin) of these events could potentially be detected by space-borne detectors, including LISA, Taiji, and TianQin. The next-generation GW detectors such as DECIGO, ET, and CE can nearly cover all these signals.

astro-ph.HE

ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian Splatting

With the rapid development of 3D reconstruction technology, the widespread distribution of 3D data has become a future trend. While traditional visual data (such as images and videos) and NeRF-based formats already have mature techniques for copyright protection, steganographic techniques for the emerging 3D Gaussian Splatting (3D-GS) format have yet to be fully explored. To address this, we propose ConcealGS, an innovative method for embedding implicit information into 3D-GS. By introducing the knowledge distillation and gradient optimization strategy based on 3D-GS, ConcealGS overcomes the limitations of NeRF-based models and enhances the robustness of implicit information and the quality of 3D reconstruction. We evaluate ConcealGS in various potential application scenarios, and experimental results have demonstrated that ConcealGS not only successfully recovers implicit information but also has almost no impact on rendering quality, providing a new approach for embedding invisible and recoverable information into 3D models in the future.

cs.CV

IIMedGPT: Promoting Large Language Model Capabilities of Medical Tasks by Efficient Human Preference Alignment

Recent researches of large language models(LLM), which is pre-trained on massive general-purpose corpora, have achieved breakthroughs in responding human queries. However, these methods face challenges including limited data insufficiency to support extensive pre-training and can not align responses with users' instructions. To address these issues, we introduce a medical instruction dataset, CMedINS, containing six medical instructions derived from actual medical tasks, which effectively fine-tunes LLM in conjunction with other data. Subsequently, We launch our medical model, IIMedGPT, employing an efficient preference alignment method, Direct preference Optimization(DPO). The results show that our final model outperforms existing medical models in medical dialogue.Datsets, Code and model checkpoints will be released upon acceptance.

cs.CL

Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding

Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high fidelity in zero-shot video tasks remains challenging. Traditional video processing methods rely heavily on fine-tuning to capture nuanced spatial-temporal details, which incurs significant data and computation costs. In contrast, training-free approaches, though efficient, often lack robustness in preserving context-rich features across complex video content. To this end, we propose DYTO, a novel dynamic token merging framework for zero-shot video understanding that adaptively optimizes token efficiency while preserving crucial scene details. DYTO integrates a hierarchical frame selection and a bipartite token merging strategy to dynamically cluster key frames and selectively compress token sequences, striking a balance between computational efficiency with semantic richness. Extensive experiments across multiple benchmarks demonstrate the effectiveness of DYTO, achieving superior performance compared to both fine-tuned and training-free methods and setting a new state-of-the-art for zero-shot video understanding.

cs.CV

RankCLIP: Ranking-Consistent Language-Image Pretraining

Self-supervised contrastive learning models, such as CLIP, have set new benchmarks for vision-language models in many downstream tasks. However, their dependency on rigid one-to-one mappings overlooks the complex and often multifaceted relationships between and within texts and images. To this end, we introduce RankCLIP, a novel pre-training method that extends beyond the rigid one-to-one matching framework of CLIP and its variants. By extending the traditional pair-wise loss to list-wise, and leveraging both in-modal and cross-modal ranking consistency, RankCLIP improves the alignment process, enabling it to capture the nuanced many-to-many relationships between and within each modality. Through comprehensive experiments, we demonstrate the effectiveness of RankCLIP in various downstream tasks, notably achieving significant gains in zero-shot classifications over state-of-the-art methods, underscoring the importance of this enhanced learning process.

cs.CV

Information Bottleneck Revisited: Posterior Probability Perspective with Optimal Transport

Information bottleneck (IB) is a paradigm to extract information in one target random variable from another relevant random variable, which has aroused great interest due to its potential to explain deep neural networks in terms of information compression and prediction. Despite its great importance, finding the optimal bottleneck variable involves a difficult nonconvex optimization problem due to the nonconvexity of mutual information constraint. The Blahut-Arimoto algorithm and its variants provide an approach by considering its Lagrangian with fixed Lagrange multiplier. However, only the strictly concave IB curve can be fully obtained by the BA algorithm, which strongly limits its application in machine learning and related fields, as strict concavity cannot be guaranteed in those problems. To overcome the above difficulty, we derive an entropy regularized optimal transport (OT) model for IB problem from a posterior probability perspective. Correspondingly, we use the alternating optimization procedure and generalize the Sinkhorn algorithm to solve the above OT model. The effectiveness and efficiency of our approach are demonstrated via numerical experiments.

cs.IT