SearcharxivSearch

arXiv subjects

Yue Wu

Publications and source records attributed to Yue Wu.

At least 19 recordsLinked to original sources

Object-Aware Background-Controlled Editing via Weighted Velocity Guidance

Training-free image editing steers diffusion or flow-matching generative models at inference time by modifying prompt-conditioned denoising velocities. Existing velocity-based editors often apply prompt-induced residuals globally over the latent space and rely on the model to localize semantic changes implicitly. For object-centric edits, these residuals are rarely zero outside the target object, so small non-target components can accumulate during multi-step integration, causing background drift and unstable object boundaries. We propose Object-Aware Velocity Control (OAVC), a training-free framework that introduces object-level control into the velocity-integration process. OAVC decouples where semantic residuals are allowed to act from how they are injected into the dynamics. It constructs a background-anchored reference interface under the source prompt and then performs object-localized safe semantic injection under the target prompt. A constrained injection operator suppresses drift-inducing velocity components, while time-adaptive spatial weighting stabilizes the transition near object boundaries. OAVC requires no training or modification of pretrained model parameters. Experiments on object-centric image and video benchmarks with image and video rectified-flow backbones show improved background preservation, structural fidelity, boundary stability, and temporal consistency while retaining effective localized editability.

cs.CV

DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection

Visual instruction tuning is crucial for advancing the vision-language alignment and instruction-following capabilities of Vision-Language Models (VLMs). However, identifying optimal subsets under a fixed ratio constraint from rapidly expanding datasets remains a significant bottleneck. While existing methods largely depend on distribution diversity or heuristic filtering, they often overlook the internal coherence within individual samples. To bridge this gap, we propose Data Intrinsic Consistency (DIC), a self-scoring metric designed to quantify the sample-level inter-component consistency. DIC consists of two modules: Visual Information Consistency (VIC), evaluating the alignment between visual content and instructions, and Response Information Consistency (RIC), assessing response coherence relative to the instruction. Building upon DIC, we introduce Data Intrinsic Consistency Selection (DICS), an adaptive data selection method that optimizes the trade-off between high intra-sample consistency and global distributional diversity under varying data budgets. Extensive experiments demonstrate that DICS consistently outperforms state-of-the-art methods across diverse dataset scales and model architectures, surpassing full-dataset fine-tuning while using only 25% of the LLaVA-1.5-665K data. We further curate DICS-6M, a 6M-sample multi-modal instruction corpus that enables the largest-scale visual instruction selection study to date; remarkably, DICS reaches 94.52\% of the official InternVL3-8B-Instruct performance using less than 25\% of its reported training data. Code can be seen at https://github.com/cqu-student/DICS

cs.CV

Influence of twist direction and large deformation on soft material torsional contact

Shear-induced contact area reduction is widely observed in soft contacts, yet recent torsional experiments have revealed a more complex non-monotonic evolution in which the contact area first increases and then decreases with twist angle. The mechanism responsible for this initial area increase and the role of large deformation in the overall area evolution remain unclear. In this study, we experimentally investigate the torsional contact response of soft Polydimethylsiloxane (PDMS) spheres by combining forward-backward twist tests with a systematic variation of the curing-agent-to-base ratio to tune material softness and deformation level. The loading-unloading tests show that the torsional interface is strongly irreversible: during unloading, the contact area follows a decrease-increase-decrease path rather than retracing the loading branch, and repeatable petal-like edges appear, indicating a wrinkle-induced surface instability. By decreasing the mixing ratio, we find that larger deformation strengthens the area-reduction contribution and eventually suppresses the initial area increase, leading to a monotonic area decrease during loading for sufficiently soft PDMS. Softer PDMS also exhibits lower shear strength, weaker torque oscillations, and improved repeatability. The results provide experimental evidence that large deformation can drive shear-induced contact area reduction, while the origin of the initial area increase remains unresolved. These findings narrow the possible mechanisms (e.g., triboelectrification) responsible for the initial area increase and provide a stringent benchmark for frictional contact models of soft interfaces.

cond-mat.soft

Intern-S2-Preview: Scientific Agentic Foundation Model

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.

cs.LG

Distribution and Transport of Fragmenting Microplastics in a 3D Global Eulerian Model

Fragmentation, the breakage of matter into smaller pieces, is an important mechanism responsible for generating microplastics (MPs). We present the first global three-dimensional Eulerian model that resolves fragmentation alongside MP transport. The evolution of particle size is modeled as a transfer from larger- to smaller-size bins, governed by a fragmentation kinetics framework. Relative to a reference simulation without fragmentation, two distinct effects are identified: (1) the surface concentration field of MPs becomes horizontally dispersed, and (2) MPs sink to depths of 500 m where the reference simulation shows negligible concentration. The vertical shift can be explained by the loss of buoyancy when particle size decreases, which facilitates horizontal sub-mixed layer transport once the particles sink below 100 m depth. Neutrally buoyant particles (with diameter d < 1 um) are continuously produced in the ocean by the fragmentation of larger particles and accumulate in the major oceanic gyres. Ultimately, the concentration of these neutrally buoyant MPs peaks at the gyre centers, a behavior that is not captured by prior models. Furthermore, the globally integrated size spectrum exhibits a steepening power-law slope over time that continues to evolve throughout our 25-year simulation. Comparisons with the AOMI Level-3wm observational dataset demonstrate a meaningful improvement in predictive skill relative to previous models: including fragmentation elevates the spatial correlation between modeled and observed surface concentrations from 45% to 58%.

physics.ao-ph

Blow-up asymptotics for a critical Hartree-type Br\'{e}zis--Nirenberg problem in dimension three

In this paper, we study the following critical Hartree problem \begin{equation}\label{equationabstract} \begin{cases} \displaystyle-\Delta u+(Q+\varepsilon V)u= A_{3,\mu} \left(\int_{\Omega}\frac{u^{6-\mu}(y)}{|x-y|^{\mu}}dy\right) u^{5-\mu} &\mathrm{~in~}\Omega,\\ \displaystyle u>0&\mathrm{~in~}\Omega,\\u=0&\mathrm{~on~}\partial\Omega,\end{cases} \end{equation} where $\Omega\subset\mathbb{R}^3$ is a bounded open set, $0<\mu<2$, the exponent $6-\mu$ is the upper critical exponent in the sense of the Hardy--Littlewood--Sobolev inequality and $A_{3,\mu}>0$ is a normalization constant. The function $Q$ is assumed to be critical in the sense of Hebey and Vaugon, and the solutions $u_{\varepsilon}$ of \eqref{equationabstract} are assumed to be an optimizing sequence for the Hardy--Littlewood--Sobolev inequality. Under a natural nondegeneracy assumption, we derive a precise asymptotic expansion of \(u_{\varepsilon}\), determine the exact blow-up rate, and identify the concentration point. We also obtain the pointwise blow-up behavior both near and away from the concentration point.

math.AP

Nonnegative Low-Rank Matrix Correction under an Orthogonality Constraint in Conservative Vlasov Simulations

In low-rank numerical methods for Vlasov dynamics, the SVD-type truncation procedure may introduce negative entries into the numerical solution. Such negative values are unphysical because the solution is a probability distribution function. We design optimization-based post-processing algorithms to recover nonnegativity while preserving the macroscopic quantities (density, momentum, and energy) pointwise. The preservation of the macroscopic quantities is written as an orthogonality constraint on the correction term. For a convex formulation based on squared nuclear norm minimization, we show that the proximal operator with the orthogonality constraint is characterized by an implicit singular value thresholding equation, and the threshold can be computed efficiently by bisection. Based on this result, we develop five algorithms for the convex formulation: Douglas--Rachford splitting, restarted dual FISTA, restarted dual accelerated gradient descent, dual PR+ conjugate gradient, and dual L-BFGS. We also consider a non-convex formulation with an explicit rank constraint and develop a tangent-space accelerated alternating projection algorithm that only requires a \(2r \times 2r\) SVD per iteration. Numerical results for a Landau damping test case show that the proposed algorithms give comparable correction quality. Among them, the tangent-space accelerated alternating projection is the most cost-efficient, increasingly so as the problem size grows. We further demonstrate the correction as a positivity limiter inside a time-dependent conservative low-rank Vlasov solver, where it removes the negativity introduced by the SVD-type truncation while preserving the conserved mass, momentum, and energy.

math.NA

PrefReward: Learning User Preference Matrix for Personalized Text Generation

Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by leveraging user histories and contextual cues. However, most existing personalization approaches rely on implicit representations within model parameters, making it difficult to interpret user-specific preferences or effectively handle long-context dependencies. To address these challenges, we propose PrefReward, a novel preference-aware generative framework that explicitly models user styles through a structured preference matrix and integrates it into the decoding process as a reward signal. PrefReward consists of two stages: (1) extracting a user-specific preference matrix that summarizes individual stylistic tendencies, and (2) using the matrix to guide generation via a KL-divergence-based reward function. Experiments on the LongLaMP dataset show that PrefReward outperforms non-personalized and retrieval-based baselines in both generation quality and personalization interpretability.

cs.CL

Probe-Conditioned Memory for Actuator-Deadband-Aware Koopman MPC in Industrial Sealing

Industrial sealing and dispensing cells often reuse a pressure chain, nozzle, substrate path, and vision interface across product recipes. For a narrow bead recipe, however, a calibrated static pressure can remain correct while small corrective moves are absorbed by actuator deadband; delivered pressure changes only after a direction- and history-dependent threshold is crossed. Commissioning is defined here as the target setup and retuning interval after such a recipe change. A physical gluing and dispensing cell provides pressure-to-width calibration, a fixed probing sequence, signal-interface limits, residual scales, and actuator bounds. The controller comparison is then run on an anonymized digital twin calibrated from those measurements. The actuator-deadband-aware Koopman model predictive controller (AK-MPC) initializes from probe-conditioned memory (PCM) that links the pressure setpoint to probe-inferred actuator behavior, a predictor, a controller prior, and a fallback filter. During commissioning, a sixteen-move probe selects a nearby historical case, fits the current pressure-width relation, updates a small local dynamic correction, and supplies a feasible receding-horizon pressure policy. In the main \(1.00\) mm benchmark, where delivered-pressure loss is visible in the probe, AK-MPC reaches 0.0487 mm tracking mean absolute error (MAE) over 60 paired cases; the calibration-only inverse, adaptive proportional-integral, online recursive-least-squares ARX, and probe-fitted ARX controllers range from 0.2492 to 0.3956 mm. This large gap reflects the full constrained Koopman-MPC and online-correction workflow. The isolated PCM contribution is measured by ablation: removing PCM raises the error to 0.0655 mm. In this regime, a short actuator characterization makes historical runs useful before much target data are available.

eess.SY

Input-to-State Stability Certification via Projection Residuals for Koopman Learning Control of Nonlinear Repetitive Systems

This paper studies input-to-state stability (ISS) certification for data-driven Koopman learning control of unknown discrete-time nonlinear repetitive systems over finite trial horizons. Rather than proposing a new learning law, we certify when a fixed Koopman-assisted constrained update yields practical stability of the selected tracking error along the trial axis. Prediction accuracy alone is insufficient for this purpose: the selected finite-horizon input-output channel must have a positive margin, and the unreachable component of the requested output increment must be accounted for through a projection residual. Thus, a Koopman predictor with small held-out prediction residuals may still fail the learning-stability certificate if its selected channel is weak. We formulate the selected stacked tracking error as the state of a discrete-time learning-axis system and treat Koopman residuals, reset mismatch, channel uncertainty, projection residuals, deployment shifts, and numerical tolerances as ISS inputs. The deterministic result gives a practical ISS estimate from the initial learning error to an explicit ultimate band. A finite-sample implementation constructs an episode-level residual bound under a fixed controller and combines it with reported channel, projection, shift, and numerical margins. Numerical checks on nonlinear repetitive systems support the predicted residual-to-band scaling, weak-channel rejection, projection closure, and ultimate-band coverage.

eess.SY

Exact Hilbert-space ergodicity from continuous monitoring

Quantum evolution is generally expected to drive a quantum many-body system toward equilibrium. This expectation is often justified by the Hilbert-space ergodicity of generic quantum dynamics, namely, the idea that pure-state evolution explores Hilbert space uniformly up to physical constraints. Such a statement can be made rigorous by requiring the associated state ensemble to form the Haar-random ensemble, or its more structured generalization, the Scrooge ensemble. In this Letter, we report the emergence of exact Hilbert-space ergodicity in a continuously monitored quantum many-body system. For any target density matrix $\sigma$, we construct a continuously monitored system for which we rigorously prove that the Scrooge ensemble of $\sigma$ is the unique late-time equilibrium distribution of quantum trajectories. Remarkably, this requires only that the jump operators in the monitoring form a deformed unitary 1-design, a seemingly much weaker condition than full ergodicity. We numerically demonstrate our predictions by simulating continuously monitored systems whose equilibrium states are thermal states. Our results establish a rigorous mechanism for the emergence of Hilbert-space ergodicity and provide a practical route for its investigation on quantum devices.

quant-ph

EP251023a: A fast X-ray transient featuring a magnetar-powered optical internal plateau followed by a steep decay

EP251023a is an extragalactic fast X-ray transient (eFXT) detected solely by EP without a gamma-ray counterpart. The prompt emission consists of a main emission with a duration $T_{90}=292\pm19$ s, followed by a long-lasting tail emission that persists until the observation ends at $T_0+1571$ s. With the upper limit of Konus--Wind, we derived a conservative upper limit on the isotropic gamma-ray energy $E_{\gamma,\rm{iso}}$ of $5.7 \times 10^{52}$ erg for the main emission phase. A redshift of $z = 2.232\pm0.001$ is identified from strong absorption features in the Keck spectrum, which also indicate a relatively low host-galaxy HI column density. Based on the broadband spectral energy distribution, the late-time light curves show an achromatic plateau, followed by an extremely steep decay with a slope of 3.99 after a break at about 49 ks, which is consistent with a rapidly spinning millisecond magnetar engine. Under the isotropic wind scenario, we obtain the initial period $P_0<2.27$~ms and the magnetic field strength $B_p<8.33\times10^{14}$~G for the magnetar; whereas considering a jet collimation with a typical opening angle of 0.1 rad relaxes these constraints to $P_0<32.15$~ms and $B_p<1.18\times10^{16}$~G. Together with GRB\,070707, EP251023a may represent a rare class of optical magnetar-powered internal plateaus with little external-shock contamination, unlike previous examples detected primarily in X-rays. Future discoveries of similar events will help clarify the relationship between magnetar-powered internal emission observed in the optical band and that detected only in X-rays.

astro-ph.HE

HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization

The quadratic complexity of attention poses a critical bottleneck for long-context processing, spurring interest in hybrid attention designs. Most open-source hybrid models adopt a layer-wise strategy. Yet, prior work has noted the inherent difficulty of integrating Linear Attention (LA) with Full Attention (FA), suggesting that the design space of attention hybridization remains underexplored. To probe this space, we conduct interpretability analysis and observe that layers exhibit block-wise functional similarity, while individual heads within the same layer display distinct functional specialization despite sharing input features. This head-level heterogeneity suggests that the head dimension provides a natural and principled granularity for fusing heterogeneous attention signals. Building on this insight, we introduce HydraHead, a novel architecture that hybridizes FA and LA along the head axis. HydraHead features two key innovations: (1) an interpretability-driven selection strategy that identifies retrieval-critical heads and preserves FA only for them, and (2) a scale-normalized fusion module that reconciles the distributional gap between FA and LA head outputs. By leveraging a three-stage transfer pipeline with parameter reuse and distillation, we achieve high-performance hybrid models with minimal training overhead. Under a unified training setup, HydraHead outperforms other hybrid designs in long-context tasks while maintaining strong general reasoning. With interpretability-driven head selection, it matches a 3:1 layer-wise hybrid's long-context performance at a 7:1 LA-to-FA ratio. Crucially, trained on only 15B tokens, HydraHead achieves over 69% improvement over the baseline at 512K context length, approaching Qwen3.5, a leading model of comparable size with a native context length of 256K. This highlights the significant scaling potential of head-level hybridization.

cs.CL

Long-time Behaviour of DLRA for SDEs

We study dynamical orthogonal (DO) approximations of stochastic differential equations and investigate their long-time behaviour. The DO formulation represents the solution by a low-rank decomposition and leads to a coupled system consisting of an evolution equation on the Stiefel manifold and a reduced stochastic process. We establish the well-posedness of the strong DO system and derive quantitative error estimates between the original stochastic differential equation and its low-rank approximation in the Wasserstein distance. Our main contribution is the analysis of invariant probability measures for the DO dynamics. Under suitable dissipativity, Lipschitz continuity, and non-degeneracy assumptions on the coefficients, we prove the existence of an invariant probability measure for the strong DO system. The proof combines uniform moment estimates, a Krylov--Bogoliubov argument for an associated frozen system, and a Kakutani-Fan-Glicksberg fixed-point theorem to recover the self-consistent dynamics. We further show that the induced low-rank process admits an invariant probability measure and discuss the structure of invariant measures through several illustrative examples. These results provide a rigorous foundation for the use of dynamical low-rank approximations in the approximation of long-time statistical properties of stochastic dynamical systems.

math.PR

Effect of Biofouling on Microplastic Transport in a 3-D Global Eulerian Model

Biofouling -- the occupation of microplastic (MP) surfaces by marine microbes -- alters particles' buoyancy and transport, yet its effect on the global distribution of MPs has not been well quantified. We present the first three-dimensional global Eulerian model to fully couple MP transport with biofouling, by augmenting the concentration field with an extra dimension representing the biomass attachment density on MP surfaces. This approach embeds time-dependent particle properties directly into the Eulerian concentration field, overcoming a fundamental challenge of tracking property evolution in grid-based models. Idealized simulations show that biofouling significantly reshapes the vertical distribution of MPs when two conditions are met: the particles must be sufficiently buoyant when they are clean to remain near the sea surface, and the local plankton growth rate must exceed the decay rate. In three-dimensional global simulations, biofouling substantially alters the distribution of large MPs ($\gtrsim 10$ $\mu$m): biofouled particles are transported below the mixed layer to 500 m depth, and the subtropical surface garbage patches become more dispersed with reduced peak concentrations. This dispersion is due to a subsurface transport route, where biofouled particles sink into layers with reversed current and are carried outward from the gyre centers before regaining buoyancy. Small particles ($\lesssim 1$ $\mu$m) remain unaffected as they stay effectively neutrally buoyant even when biofouled. A comparison with a global trawler dataset shows that incorporating biofouling reduces the fraction of outlying model-observation data points from 25\% to 13\%, demonstrating a meaningful improvement in model skill.

physics.flu-dyn

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts largely focus on text-dominant settings, leaving long-horizon multimodal tasks underexplored. This gap is evident in video tasks requiring sustained temporal understanding and iterative interaction. We present InternVideo3, a framework enhancing these capabilities via Multimodal Contextual Reasoning (MCR). MCR treats understanding as a closed-loop process over a shared, evolving context containing observations, instructions, reasoning, tool actions, and memory. This frames long-video understanding as evidence accumulation and verification. To ensure efficiency, we introduce Multimodal Multi-head Latent Attention (M^2LA), a token-preserving reparameterization compressing KV-cache states while retaining the full token stream. Our staged training includes continued pretraining, short-to-long supervised fine-tuning, rule-based reinforcement learning, and on-policy distillation. Experiments show InternVideo3 achieves strong performance on benchmarks like Video-MME, MLVU, and EgoSchema. We further instantiate the model as a video agent with retrieval tools, demonstrating robust evidence-grounded behavior. Our results suggest that efficient context handling and closed-loop reasoning are vital for adapting open multimodal models toward long-horizon visually grounded agency.

cs.CV

PTL-Diffusion: Manifold-Aware Diffusion with Periodic Terminal Laws

Standard diffusion models typically use a single time-homogeneous Gaussian terminal distribution as the reference law for generation. While this choice is analytically convenient and empirically powerful, it provides little explicit structure for data concentrated near low-dimensional manifolds, where different regions of the data distribution may correspond to distinct local geometric or semantic factors. As a result, the reverse model must recover manifold-level structure almost entirely from an unstructured terminal reference distribution. We propose PTL-Diffusion, a proof-of-concept diffusion framework whose forward noising process converges to a nonconstant periodic family of Gaussian terminal laws rather than to a single invariant law. Unlike a phase-conditioned DDPM, where phase information only enters the denoising network while the forward process remains unchanged, PTL-Diffusion embeds phase structure directly into the forward noising dynamics. The proposed construction remains close to standard denoising diffusion models: for a periodically forced Ornstein--Uhlenbeck-type forward process, we derive closed-form forward marginals, the limiting periodic Gaussian terminal family, and explicit Gaussian reverse posteriors, enabling standard noise-prediction training. We also introduce an invariant-average regularization term coupling the phase-conditioned reverse dynamics through the averaged periodic reference law. Experiments on torus and cylinder point-cloud benchmarks and the Olivetti face dataset show that PTL-Diffusion improves manifold-level distributional matching over matched DDPM baselines, reducing phase-conditioned errors, feature-space covariance errors, and nearest-neighbour manifold distances. These results suggest structured terminal reference laws as a promising direction, while motivating more expressive phase constructions and larger-scale evaluations.

cs.CV

MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models

Conventional Post-Training Quantization (PTQ) methods struggle with 4-bit Omni-modal Large Language Models (OLLMs) due to the extreme distribution heterogeneity and disparate outlier patterns across modalities. To address this, we propose MorphoQuant, a modality-aware PTQ framework engineered to preserve cross-modal morphology and mitigate outlier loss. Specifically, we introduce Distribution-Aware Bias Compensation (DABC), which selectively absorbs long-tailed outliers into channel-wise biases. This mechanism safeguards outlier magnitudes while maintaining high-precision discretization for dense inliers, thereby preserving accurate discretization across diverse modal distribution. Complementing this, we propose Morphology-Directed Quantization Function Optimization (MDQFO) to co-optimize the quantization grid with the bias mask, ensuring fine-grained alignment across modalities. Extensive evaluations on Qwen2.5-Omni across benchmarks like MMMU and Video-MME demonstrate our approach's superiority. Notably, our W4A4 model achieves 76.63% on ScienceQA, significantly outperforming SOTA W4A4 methods and surprisingly surpassing the W4A16 baseline, which fully demonstrates the exceptional accuracy-efficiency trade-off of our framework.

cs.CV