SearcharxivSearch

arXiv subjects

Tianyi Chu

Publications and source records attributed to Tianyi Chu.

17 recordsLinked to original sources

Coupled Rayleigh--Taylor and Faraday instabilities in vertically vibrated cylindrical containers

Interfacial instabilities govern the mixing in confined multiphase flows. Yet, the two mechanisms that drive them are usually studied independently: the pressure-gradient-driven Rayleigh--Taylor (RT) instability, which amplifies long-wavelength modes, and the parametrically forced Faraday instability, which selects shorter-wavelength harmonic or subharmonic modes. When an adverse density contrast and vertical vibration act together, the two compete, and neither alone describes the response. We use Floquet analysis to characterize the onset, growth, modal structure, and velocity fields of Faraday--RT waves in a vertically vibrated cylinder, resolved by azimuthal wavenumber, radial (Bessel) mode, and Floquet harmonic. The formulation recovers the classical RT and Faraday limits and reproduces the instability onset at the frequencies measured experimentally. For a free-sliding interface, increasing the vibration amplitude shifts the dominant instability mechanism from RT growth to subharmonic and then harmonic Faraday responses. Lateral confinement can also stabilize individual RT modes, which is not possible in an unbounded domain, although other Faraday modes may remain unstable. Pinning the contact line couples radial modes that otherwise evolve independently, allowing the unstable mode to be a superposition of RT-unstable and Faraday-stable components. This superposition alters the instability mechanism, producing a richer radial pattern. Reconstruction of the unstable modes shows the (linear) velocity fields that imaging cannot access and demonstrates how the instabilities can change the flow more broadly.

physics.flu-dyn

Limits of constant-parameter constitutive models for hydrogels under inertial cavitation

Mechanical characterization of soft materials at high strain rates is challenging due to their high compliance, nonlinear viscoelastic behavior, and potentially history-dependent responses. Inertial microcavitation rheometry (IMR) addresses this challenge by coupling laser-induced cavitation (LIC) experiments with numerical simulations of bubble dynamics models to infer constitutive models and material parameters. Both IMR and its variants infer parameters that depend on the chosen fitting window, which suggests that a constant-parameter constitutive model is insufficient to describe the full cavitation event. We use this window dependence to identify when the constant-parameter assumption fails, rather than to report a single effective parameter set. The constitutive parameters are estimated over moving, overlapping windows using a modified iterative ensemble Kalman smoother with multiple data assimilation (MIEnKS-MDA). Within the neo-Hookean Kelvin--Voigt (NHKV) constitutive model, we obtain time-resolved estimates of the constitutive response in polyacrylamide (PAAm) hydrogels with different crosslinker concentrations. The inferred shear modulus and viscosity generally decrease and then plateau during cavitation, while exhibiting relatively weak temperature sensitivity. For gelatin gels, by contrast, the inferred property evolution shows a pronounced temperature dependence, with distinct trends at low and high temperatures. Moreover, both the apparent shear modulus and viscosity exhibit significant variations during the first two bubble collapses. These results show that time-resolved parameter estimation within the prescribed NHKV constitutive structure can diagnose where the constant-parameter model assumption falls short during cavitation, thereby guiding the development of improved physics-based models of complex bubble--material interactions.

cond-mat.soft

M$^2$-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining

Graphical User Interface (GUI) agent is pivotal to advancing intelligent human-computer interaction paradigms. Constructing powerful GUI agents necessitates the large-scale annotation of high-quality user-behavior trajectory data (i.e., intent-trajectory pairs) for training. However, manual annotation methods and current GUI agent data mining approaches typically face three critical challenges: high construction cost, poor data quality, and low data richness. To address these issues, we propose M$^2$-Miner, the first low-cost and automated mobile GUI agent data-mining framework based on Monte Carlo Tree Search (MCTS). For better data mining efficiency and quality, we present a collaborative multi-agent framework, comprising InferAgent, OrchestraAgent, and JudgeAgent for guidance, acceleration, and evaluation. To further enhance the efficiency of mining and enrich intent diversity, we design an intent recycling strategy to extract extra valuable interaction trajectories. Additionally, a progressive model-in-the-loop training strategy is introduced to improve the success rate of data mining. Extensive experiments have demonstrated that the GUI agent fine-tuned using our mined data achieves state-of-the-art performance on several commonly used mobile GUI benchmarks. Our work will be released to facilitate the community research.

cs.AI

Energy dissipation mechanisms in an acoustically-driven slit

We quantify how incident acoustic energy is converted into vortical motion and viscous dissipation for a two-dimensional plane-wave passing through a slit geometry. We perform direct numerical simulations over a broad parameter space in incident sound pressure level (ISPL), Strouhal number (St), and Reynolds number (Re). Spectral proper orthogonal decomposition (SPOD) yields energy-ranked coherent structures at each frequency, from which we construct mode-by-mode fields for spectral kinetic energy (KE) and viscous loss (VL) components to examine the mechanisms of acoustic absorption. At ISPL=150dB, the acoustic-hydrodynamic energy conversion is highest when the acoustic displacement amplitude is comparable to the slit thickness, corresponding to a Keulegan-Carpenter number of order unity. In this regime, the oscillatory boundary layer undergoes periodic separation, resulting in vortex shedding that dominates acoustic damping. VL accounts for 20-60% of the KE contribution. For higher acoustic frequencies, the confinement of the Stokes layer produces X-shaped near-slit modes, reducing the total energy input by approximately 50%. The influence of Re depends on amplitude. At ISPL=150dB, larger Re values correspond to suppressed broadband fluctuations and sharpened harmonic peaks. At ISPL = 120dB, the boundary layers remain attached, vortex shedding is weak, absorption monotonically scales with viscosity, and the Re- and St-dependencies become comparable. Across all conditions, more than 99% of the VL is confined to a compact region surrounding the slit mouth. The KE-VL spectra describe parameter regimes that enhance or suppress acoustic damping in slit geometries, providing a physically interpretable basis for acoustic-based design.

physics.flu-dyn

DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation

Personalized Text-to-Image (PT2I) generation aims to produce customized images based on reference images. A prominent interest pertains to the integration of an image prompt adapter to facilitate zero-shot PT2I without test-time fine-tuning. However, current methods grapple with three fundamental challenges: 1. the elusive equilibrium between Concept Preservation (CP) and Prompt Following (PF), 2. the difficulty in retaining fine-grained concept details in reference images, and 3. the restricted scalability to extend to multi-subject personalization. To tackle these challenges, we present Dynamic Image Prompt Adapter (DynaIP), a cutting-edge plugin to enhance the fine-grained concept fidelity, CP-PF balance, and subject scalability of SOTA T2I multimodal diffusion transformers (MM-DiT) for PT2I generation. Our key finding is that MM-DiT inherently exhibit decoupling learning behavior when injecting reference image features into its dual branches via cross attentions. Based on this, we design an innovative Dynamic Decoupling Strategy that removes the interference of concept-agnostic information during inference, significantly enhancing the CP-PF balance and further bolstering the scalability of multi-subject compositions. Moreover, we identify the visual encoder as a key factor affecting fine-grained CP and reveal that the hierarchical features of commonly used CLIP can capture visual information at diverse granularity levels. Therefore, we introduce a novel Hierarchical Mixture-of-Experts Feature Fusion Module to fully leverage the hierarchical features of CLIP, remarkably elevating the fine-grained concept fidelity while also providing flexible control of visual granularity. Extensive experiments across single- and multi-subject PT2I tasks verify that our DynaIP outperforms existing approaches, marking a notable advancement in the field of PT2l generation.

cs.CV

SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing

The existing Multimodal Large Language Models (MLLMs) for GUI perception have made great progress. However, the following challenges still exist in prior methods: 1) They model discrete coordinates based on text autoregressive mechanism, which results in lower grounding accuracy and slower inference speed. 2) They can only locate predefined sets of elements and are not capable of parsing the entire interface, which hampers the broad application and support for downstream tasks. To address the above issues, we propose SparkUI-Parser, a novel end-to-end framework where higher localization precision and fine-grained parsing capability of the entire interface are simultaneously achieved. Specifically, instead of using probability-based discrete modeling, we perform continuous modeling of coordinates based on a pre-trained Multimodal Large Language Model (MLLM) with an additional token router and coordinate decoder. This effectively mitigates the limitations inherent in the discrete output characteristics and the token-by-token generation process of MLLMs, consequently boosting both the accuracy and the inference speed. To further enhance robustness, a rejection mechanism based on a modified Hungarian matching algorithm is introduced, which empowers the model to identify and reject non-existent elements, thereby reducing false positives. Moreover, we present ScreenParse, a rigorously constructed benchmark to systematically assess structural perception capabilities of GUI models across diverse scenarios. Extensive experiments demonstrate that our approach consistently outperforms SOTA methods on ScreenSpot, ScreenSpot-v2, CAGUI-Grounding and ScreenParse benchmarks. The resources are available at https://github.com/antgroup/SparkUI-Parser.

cs.AI

Competing Mechanisms at Vibrated Interfaces of Density-Contrast Fluids

Fluid--fluid interfacial instability and subsequent fluid mixing are ubiquitous in nature and engineering. The hydrodynamic instability of fluid interfaces has long centered on the pressure gradient-driven long-wavelength Rayleigh--Taylor instability and the resonance-induced short-wavelength Faraday instability. However, neither instability alone can explain the dynamics when both mechanisms are present. We identify a previously unseen multi-modal instability emerging from their coexistence. When the denser fluid is polydimethylsiloxane, the mixed region at a high density contrast (Atwood number=0.9) spans a vibration amplitude range approximately twice the gravitational acceleration. Using Floquet stability analysis, we show how vibrations govern transitions between the RT and Faraday instabilities, leading to contention between these instabilities rather than resonant enhancement. The initial transient growth is represented by the exponential modal growth of the most unstable Floquet exponent, along with its accompanying periodic behavior. Direct numerical simulations validate these findings and track interface breakup into the multiscale and nonlinear regimes. Specifically, we show that growing RT modes nonlinearly suppress Faraday responses even when the initial growth rate of the Faraday instability is 3.63 times that of RT, so a bidirectional competition hinders their sustained coexistence.

physics.flu-dyn

Accelerating Bayesian Optimal Experimental Design via Local Radial Basis Functions: Application to Soft Material Characterization

We develop a computational approach that significantly improves the efficiency of Bayesian optimal experimental design (BOED) using local radial basis functions (RBFs). The presented RBF--BOED method uses the intrinsic ability of RBFs to handle scattered parameter points, a property that aligns naturally with the probabilistic sampling inherent in Bayesian methods. By constructing accurate deterministic surrogates from local neighborhood information, the method enables high-order approximations with reduced computational overhead. As a result, computing the expected information gain (EIG) requires evaluating only a small uniformly sampled subset of prior parameter values, greatly reducing the number of expensive forward-model simulations needed. For demonstration, we apply RBF--BOED to optimize a laser-induced cavitation (LIC) experimental setup, where forward simulations follow from inertial microcavitation rheometry (IMR) and characterize the viscoelastic properties of hydrogels. Two experimental design scenarios, single- and multi-constitutive-model problems, are explored. Results show that EIG estimates can be obtained at just 8% of the full computational cost in a five-model problem within a two-dimensional design space. This advance offers a scalable path toward optimal experimental design in soft and biological materials.

physics.comp-ph

Stochastic reduced-order Koopman model for turbulent flows

A stochastic data-driven reduced-order model applicable to a wide range of turbulent natural and engineering flows is presented. Combining ideas from Koopman theory and spectral model order reduction, the stochastic low-dimensional inflated convolutional Koopman model (SLICK) accurately forecasts short-time transient dynamics while preserving long-term statistical properties. A discrete Koopman operator is used to evolve convolutional coordinates that govern the temporal dynamics of spectral orthogonal modes, which in turn represent the energetically most salient large-scale coherent flow structures. Turbulence closure is achieved in two steps: first, by inflating the convolutional coordinates to incorporate nonlinear interactions between different scales, and second, by modeling the residual error as a stochastic source. An empirical dewhitening filter informed by the data is used to maintain the second-order flow statistics within the long-time limit. The model uncertainty is quantified through either Monte Carlo simulation or by directly propagating the model covariance matrix. The model is demonstrated on the Ginzburg-Landau equations, large-eddy simulation (LES) data of a turbulent jet, and particle image velocimetry (PIV) data of the flow over an open cavity. In all cases, the model is predictive over time horizons indicated by a detailed error analysis and integrates stably over arbitrary time horizons, generating realistic surrogate data.

physics.flu-dyn

Revealing Structure and Symmetry of Nonlinearity in Natural and Engineering Flows

Energy transfer across scales is fundamental in fluid dynamics, linking large-scale flow motions to small-scale turbulent structures in engineering and natural environments. Triadic interactions among three wave components form complex networks across scales, challenging understanding and model reduction. We introduce Triadic Orthogonal Decomposition (TOD), a method that identifies coherent flow structures optimally capturing spectral momentum transfer, quantifies their coupling and energy exchange in an energy budget bispectrum, and reveals the regions where they interact. TOD distinguishes three components--a momentum recipient, donor, and catalyst--and recovers laws governing pairwise, six-triad, and global triad conservation. Applied to unsteady cylinder wake and wind turbine wake data, TOD reveals networks of triadic interactions with forward and backward energy transfer across frequencies and scales.

physics.flu-dyn

Bayesian optimal design accelerates discovery of material properties from bubble dynamics

An optimal sequential experimental design approach is developed to computationally characterize soft material properties at the high strain rates associated with bubble cavitation. The approach involves optimal design and model inference. The optimal design strategy maximizes the expected information gain in a Bayesian statistical setting to design experiments that provide the most informative cavitation data about unknown soft material properties. We infer constitutive models by characterizing the associated viscoelastic properties from measurements via a hybrid ensemble-based 4D-Var method (En4D-Var). The inertial microcavitation-based high strain-rate rheometry (IMR) method ([1]) simulates the bubble dynamics under laser-induced cavitation. We use experimental measurements to create synthetic data representing the viscoelastic behavior of stiff and soft polyacrylamide hydrogels under realistic uncertainties. The synthetic data are seeded with larger errors than state-of-the-art measurements yet match known material properties, reaching 1% relative error within 10 sequential designs (experiments). We discern between two seemingly equally plausible constitutive models, Neo-Hookean Kelvin--Voigt and quadratic Kelvin--Voigt, with a probability of correctness larger than 99% in the same number of experiments. This strategy discovers soft material properties, including discriminating between constitutive models and discerning their parameters, using only a few experiments.

cond-mat.soft

PNeSM: Arbitrary 3D Scene Stylization via Prompt-Based Neural Style Mapping

3D scene stylization refers to transform the appearance of a 3D scene to match a given style image, ensuring that images rendered from different viewpoints exhibit the same style as the given style image, while maintaining the 3D consistency of the stylized scene. Several existing methods have obtained impressive results in stylizing 3D scenes. However, the models proposed by these methods need to be re-trained when applied to a new scene. In other words, their models are coupled with a specific scene and cannot adapt to arbitrary other scenes. To address this issue, we propose a novel 3D scene stylization framework to transfer an arbitrary style to an arbitrary scene, without any style-related or scene-related re-training. Concretely, we first map the appearance of the 3D scene into a 2D style pattern space, which realizes complete disentanglement of the geometry and appearance of the 3D scene and makes our model be generalized to arbitrary 3D scenes. Then we stylize the appearance of the 3D scene in the 2D style pattern space via a prompt-based 2D stylization algorithm. Experimental results demonstrate that our proposed framework is superior to SOTA methods in both visual quality and generalization.

cs.CV

Attack Deterministic Conditional Image Generative Models for Diverse and Controllable Generation

Existing generative adversarial network (GAN) based conditional image generative models typically produce fixed output for the same conditional input, which is unreasonable for highly subjective tasks, such as large-mask image inpainting or style transfer. On the other hand, GAN-based diverse image generative methods require retraining/fine-tuning the network or designing complex noise injection functions, which is computationally expensive, task-specific, or struggle to generate high-quality results. Given that many deterministic conditional image generative models have been able to produce high-quality yet fixed results, we raise an intriguing question: is it possible for pre-trained deterministic conditional image generative models to generate diverse results without changing network structures or parameters? To answer this question, we re-examine the conditional image generation tasks from the perspective of adversarial attack and propose a simple and efficient plug-in projected gradient descent (PGD) like method for diverse and controllable image generation. The key idea is attacking the pre-trained deterministic generative models by adding a micro perturbation to the input condition. In this way, diverse results can be generated without any adjustment of network structures or fine-tuning of the pre-trained models. In addition, we can also control the diverse results to be generated by specifying the attack direction according to a reference text or image. Our work opens the door to applying adversarial attack to low-level vision tasks, and experiments on various conditional image generation tasks demonstrate the effectiveness and superiority of the proposed method.

cs.CV

Mesh-Free Hydrodynamic Stability

A specialized mesh-free radial basis function-based finite difference (RBF-FD) discretization is used to solve the large eigenvalue problems arising in hydrodynamic stability analyses of flows in complex domains. Polyharmonic spline functions with polynomial augmentation (PHS+poly) are used to construct the discrete linearized incompressible and compressible Navier-Stokes operators on scattered nodes. Rigorous global and local eigenvalue stability studies of these global operators and their constituent RBF stencils provide a set of parameters that guarantee stability while balancing accuracy and computational efficiency. Specialized elliptical stencils to compute boundary-normal derivatives are introduced and the treatment of the pole singularity in cylindrical coordinates is discussed. The numerical framework is demonstrated and validated on a number of hydrodynamic stability methods ranging from classical linear theory of laminar flows to state-of-the-art non-modal approaches that are applicable to turbulent mean flows. The examples include linear stability, resolvent, and wavemaker analyses of cylinder flow at Reynolds numbers ranging from 47 to 180, and resolvent and wavemaker analyses of the self-similar flat-plate boundary layer at a Reynolds number as well as the turbulent mean of a high-Reynolds-number transonic jet at Mach number 0.9. All previously-known results are found in close agreement with the literature. Finally, the resolvent-based wavemaker analyses of the Blasius boundary layer and turbulent jet flows offer new physical insight into the modal and non-modal growth in these flows.

physics.flu-dyn

Joint Optimization of Triangle Mesh, Material, and Light from Neural Fields with Neural Radiance Cache

Traditional inverse rendering techniques are based on textured meshes, which naturally adapts to modern graphics pipelines, but costly differentiable multi-bounce Monte Carlo (MC) ray tracing poses challenges for modeling global illumination. Recently, neural fields has demonstrated impressive reconstruction quality but falls short in modeling indirect illumination. In this paper, we introduce a simple yet efficient inverse rendering framework that combines the strengths of both methods. Specifically, given pre-trained neural field representing the scene, we can obtain an initial estimate of the signed distance field (SDF) and create a Neural Radiance Cache (NRC), an enhancement over the traditional radiance cache used in real-time rendering. By using the former to initialize differentiable marching tetrahedrons (DMTet) and the latter to model indirect illumination, we can compute the global illumination via single-bounce differentiable MC ray tracing and jointly optimize the geometry, material, and light through back propagation. Experiments demonstrate that, compared to previous methods, our approach effectively prevents indirect illumination effects from being baked into materials, thus obtaining the high-quality reconstruction of triangle mesh, Physically-Based (PBR) materials, and High Dynamic Range (HDR) light probe.

cs.GR

RBF-FD discretization of the Navier-Stokes equations on scattered but staggered nodes

A semi-implicit fractional-step method that uses a staggered node layout and radial basis function-finite differences (RBF-FD) to solve the incompressible Navier-Stokes equations is developed. Polyharmonic splines (PHS) with polynomial augmentation (PHS+poly) are used to construct the global differentiation matrices. A systematic parameter study identifies a combination of stencil size, PHS exponent, and polynomial degree that minimizes the truncation error for a wave-like test function on scattered nodes. Classical modified wavenumber analysis is extended to RBF-FDs on heterogeneous node distributions and used to confirm that the accuracy of the selected 28-point stencil is comparable to that of spectral-like, 6th-order Padé-type finite differences. The Navier-Stokes solver is demonstrated on two benchmark problems, internal flow in a lid-driven cavity in the Reynolds number regime $10^2\leq$Re$\leq10^4$, and open flow around a cylinder at Re=100 and 200. The combination of grid staggering and careful parameter selection facilitates accurate and stable simulations at significantly lower resolutions than previously reported, using more compact RBF-FD stencils, without special treatment near solid walls, and without the need for hyperviscosity or other means of regularization.

physics.flu-dyn

A stochastic SPOD-Galerkin model for broadband turbulent flows

The use of spectral proper orthogonal decomposition (SPOD) to construct low-order models for broadband turbulent flows is explored. The choice of SPOD modes as basis vectors is motivated by their optimality and space-time coherence properties for statistically stationary flows. This work follows the modeling paradigm that complex nonlinear fluid dynamics can be approximated as stochastically forced linear systems. The proposed stochastic two-level SPOD-Galerkin model governs a compound state consisting of the modal expansion coefficients and forcing coefficients. In the first level, the modal expansion coefficients are advanced by the forced linearized Navier-Stokes operator under the linear time-invariant assumption. The second level governs the forcing coefficients, which compensate for the offset between the linear approximation and the true state. At this level, least squares regression is used to achieve closure by modeling nonlinear interactions between modes. The statistics of the remaining residue are used to construct a dewhitening filter that facilitates the use of white noise to drive the model. If the data residue is used as the sole input, the model accurately recovers the original flow trajectory for all times. If the residue is modeled as stochastic input, then the model generates surrogate data that accurately reproduces the second-order statistics and dynamics of the original data. The stochastic model uncertainty, predictability, and stability are quantified analytically and through Monte Carlo simulations. The model is demonstrated on large eddy simulation data of a turbulent jet at Mach number $M=0.9$ and Reynolds number of $Re_D\approx 10^6$.

physics.flu-dyn