SearcharxivSearch

arXiv subjects

Zhan Zhang

Publications and source records attributed to Zhan Zhang.

At least 19 recordsLinked to original sources

Flexible Motion Generation from Language and Style References

We introduce FlexMoGen, a novel framework for flexible human motion synthesis conditioned on both natural language descriptions and motion style references. Text prompts are effective at defining semantic content, but they are often limited in capturing fine-grained style details such as timing, limb articulation, and expressive dynamics. A style example clip supplements the text by conveying these nuanced motion characteristics directly, enabling the model to preserve high-level intent while reproducing the desired stylistic traits. Given a text prompt and a style example clip, FlexMoGen generates high-quality motions that preserve semantic content while faithfully reflecting the target style, offering users greater control over the animation generation process. Unlike prior methods that rely on discrete style labels and do not generalize to long or multi-style generation, FlexMoGen learns a variational style encoder without style supervision and supports long, time-varying, multi-style synthesis. Our framework jointly pre-trains the style encoder and a text-to-motion latent diffusion model within a unified architecture, modulating motion style through a lightweight adaptation module. It integrates an efficient relative positional encoding scheme and is trained on both stylized and non-stylized datasets, enabling strong generalization to unseen text-style combinations. Experiments show that FlexMoGen achieves the best balance between content fidelity and style reflection.

cs.CV

VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows

Accurate prediction of unsteady separated flows is challenging because the aerodynamic loads depend on nonlinear separation and vortex-shedding dynamics. Although high-fidelity CFD resolves these mechanisms, its cost limits repeated use in design and control. Standard field-level surrogate training, however, does not distinguish the flow regions that contribute most strongly to the aerodynamic loads. We introduce VATO (Vortex-Force-Aware Transformer Operator), which couples the Vortex Force Map (VFM) method to a geometry-aware neural operator through two complementary mechanisms. VATO-S adds training-only supervision of the local VFM force-contribution field, with no increase in model size or inference cost. VATO-A uses VFM contribution and sensitivity fields to prioritise force-relevant source locations for residual cross attention. The methods are evaluated on unsteady CFD data for double-edged-plate aerofoils over 54 trajectories from nine geometries. Over lead times of 1-20~ms, VATO-S reduces velocity, pressure, and vorticity errors by 10.4\%, 1.0\%, and 15.6\%, respectively, while VATO-A achieves reductions of 15.8\%, 7.5\%, and 31.2\%. VATO-S gives the lowest VFM-derived drag error, whereas VATO-A gives the lowest pressure-derived lift and drag errors. Over lead times extending 50\% beyond the training range, VATO-A retains a 26.9\% reduction in vorticity error and larger improvements in all four force readouts, despite reduced gains in velocity and pressure. These results show that force-aware operator learning can improve both flow-field prediction and aerodynamic functional accuracy in unsteady separated flows.

cs.LG

Charge transfer and competing symmetry breaking drive orbital reconstruction and emergent ferromagnetism in insulating oxide superlattices

Electron correlation, hopping, and ligand-to-metal charge transfer collectively lead to diverse electronic and magnetic phenomena in 3$d$ transition-metal oxides, where directional d orbitals make hopping highly sensitive to symmetry-dependent orbital overlap. Heterostructure engineering with atomically flat interfaces adds symmetry-breaking charge transfer as a further route to emergent behavior, yet whether interfacial mismatch between constituent oxides of a superlattice shapes ground states independent of epitaxial strain remains unresolved. Here we examine superlattices combining NdNiO$_3$ with Mott-insulating NdMnO$_3$. Varying layer thickness and combining transport with X-ray spectroscopy, we show that electron transfer from NdMnO$_3$ to NdNiO$_3$ drives a room-temperature insulating state with a distinct electronic structure, accompanied by a reversal in orbital symmetry beyond simple strain considerations, underscoring the interface's central role. These reconstructions stabilize an emergent ferromagnetic insulating state arising from interfacial Ni$^{2+}$-O-Mn$^{4+}$ superexchange. Our results establish a pathway to interface-engineered ferromagnetic insulating phases via competing interactions, with potential for spin-insulatronic applications.

cond-mat.mtrl-sci

STCO: Conditional Neural Operators for Time-Dependent PDEs

Neural operators have emerged as efficient surrogates for time-dependent physical systems governed by partial differential equations (PDEs), but their future-state predictions are often conditioned only on observed states and static problem descriptors. For control or optimization, however, body motion, inflow, or forcing are prescribed for the query without being determined solely by the observed state. We introduce the Spatiotemporal Conditional Operator (STCO) for prescribed-condition operator learning (PCOL), a common interface that supplies prescribed target-time condition fields to heterogeneous backbone architectures while retaining their architecture-specific core computation and context pathways. Its condition interface combines Flow-Aware Graph Leaf (FAGL) with Dual-Site Feature-wise Linear Modulation (DSFiLM). Non-learned FAGL uses vorticity from the final observed frame to construct a fixed-cardinality adaptive partition, then co-locates the observed history and target-time condition fields at its regional coordinates. DSFiLM injects separate motion, inflow, and force routes before and after operator computation through current-feature-driven slot- and channel-wise gates. We evaluate twelve matched backbone architectures with different existing physical and temporal inputs. The immersed-boundary computational fluid dynamics (CFD) benchmark spans prescribed motion, inflow disturbances, body-force actuation, and morphology. Across twelve matched backbones, three regimes, and two lead ranges, STCO yields mean paired reductions of 31.1% in relative-L2 field error and 24.7% in normalized pressure-derived load error. It also lowers longer-lead field error for 11 backbones, while interventions on individual condition groups produce measurable prediction changes for every group evaluated.

cs.AI

Plume Segmentation from MethaneSAT with Cross-Sensor Transfer Learning and Physics-Informed Postprocessing

Automated detection and masking of individual methane plumes from satellite imagery is important for operational emission attribution and quantification. We present a machine learning framework for plume detection from MethaneSAT retrieved column-averaged dry-air mole fractions of methane. We address two core challenges: the scarcity of labeled MethaneSAT data and the need for inference reliability across diverse atmospheric and surface conditions. We first demonstrate that Mask R-CNN with a ResNet-50 backbone outperforms U-Net semantic segmentation on both MethaneAIR (an airborne version of MethaneSAT) and MethaneSAT data, with pixel-level F1 score gains of 10.49 and 5.48 respectively. To address MethaneSAT data scarcity, we evaluate three cross-sensor transfer strategies leveraging MethaneAIR flights and synthetic plumes. Mask R-CNN with ResNet-50 fine-tuned from MethaneAIR pre-trained weights is the most effective strategy, achieving instance-level precision of 0.60 and a near-perfect recall of 0.98 at the baseline operating point. A physics-informed post-processing pipeline converts detections into two operationally distinct modes. The first is a high-sensitivity mode that applies morphological filtering and proximity-based merging for comprehensive emission screening, achieving precision of 0.71 and recall of 0.94. The second is a high-precision mode that additionally applies a distribution-based classifier for confident source attribution, achieving precision of 0.92 and recall of 0.70. Manual review of detections classified as false positives against our wavelet-based ground truth labels reveals that a meaningful fraction of cases correspond to real methane enhancements excluded by conservative labeling criteria, indicating that precision values reported are lower bounds on true detection performance... Our data and code are available at: https://doi.org/10.7910/DVN/FR959H

cs.CV

The Reynolds-Averaged Vortex Force Map Method

Vortex-force mapping (VFM) links vortical flow structures to aerodynamic forces through compact-domain integrals weighted by geometry-only Laplace potentials, but existing formulations are tied to simple geometries and laminar flows. In this study, we derive a Reynolds-averaged vortex force map (RA-VFM) directly from the incompressible Reynolds-averaged Navier-Stokes (RANS) equations, augmenting the classical vortex-pressure (VP) term with a Reynolds-stress (RS) contribution based on the Laplace-potential-weighted divergence of the modelled Reynolds stress (Boussinesq eddy-viscosity form). The resulting framework reconstructs mean lift and drag from RANS mean fields while retaining spatial attribution of force production to specific regions and coherent structures within a compact control volume. We apply RA-VFM to unsteady RANS ($k$-$\omega$ SST) simulations of a realistic gliding goshawk with strong three-dimensionality and a matched GOE803 aerofoil section. For the aerofoil, the VP term alone reproduces the CFD force curves over the pre- and near-stall range, with RS contributions becoming appreciable only in deep stall. For the bird, by contrast, the VP term underpredicts both $C_L$ and $C_D$, whereas including the RS term reduces the mean absolute error relative to CFD from $6\%$ to $2\%$ in lift and from $5\%$ to $1\%$ in drag over an angle of attack range of $0^\circ$-$20^\circ$. RA-VFM thus extends vortex-force mapping to turbulent, 3-D RANS flows and enables quantitative attribution of mean lift and drag to specific coherent structures within compact domains.

physics.flu-dyn

Thermodynamic Focusing for Inference-Time Search: Practical Methods for Target-Conditioned Sampling and Prompted Inference

Finding rare but useful solutions in very large candidate spaces is a recurring practical challenge across language generation, planning, and reinforcement learning. We present a practical framework, \emph{Inverted Causality Focusing Algorithm} (ICFA), that treats search as a target-conditioned reweighting process. ICFA reuses an available proposal sampler and a task-specific similarity function to form a focused sampling distribution, while adaptively controlling focusing strength to avoid degeneracy. We provide a clear recipe, a stability diagnostic based on effective sample size, a compact theoretical sketch explaining when ICFA can reduce sample needs, and two reproducible experiments: constrained language generation and sparse-reward navigation. We further show how structured prompts instantiate an approximate, language-level form of ICFA and describe a hybrid architecture combining prompted inference with algorithmic reweighting.

cs.LG

Bergman Projections, Kernel $p$-Norm Estimates, and Toeplitz Operators with B\'{e}koll\'{e} and Bonami weights

In this paper, we establish entirely new $p$-norm estimates for reproducing kernels to characterize the bounded and compact Toeplitz operators $T_{\mu}$ acting between weighted B\'{e}koll\'{e}--Bonami Bergman spaces $A^p_u(\mathbb{D})$ and $A^q_u(\mathbb{D})$ for all positive exponents $0 < p, q < \infty$. These operator-theoretic properties are completely described in terms of generalized Berezin transforms, averaging functions, and Carleson measures. We introduce two explicit conditions on the weights to ensure the boundedness of the weighted Bergman projection $P_u$, generalizing results from Hilbert spaces to Banach spaces.Our work generalizes the main results of Tong, Li, and Arroussi \cite{TLA} from Hilbert spaces to the more general setting of Banach spaces.

math.CV

Elliptic functions, Floquet transform and Bergman spaces on doubly periodic domains

We study Bergman spaces A^2(D), their kernels and Toeplitz operators on unbounded, doubly periodic domains D in the complex plane. We establish the mapping properties of the Floquet transform operator defined in A^2(D) and derive a general formula connecting the Bergman kernel and projection of the domain D to a kernel and projection on the bounded periodic cell B. As an application, we prove, for Toeplitz operators T_a with doubly periodic symbols, a spectral band formula, which describes the spectrum and essential spectrum of T_a in terms of the spectra of a family of Toeplitz-type operators on the cell B. Technical challenges arise from the fact that double quasiperiodic boundary conditions have to be taken into account in the definitions of the spaces and operators on the periodic cell B. This requires novel operator theoretic tools, which are based on modifications of certain elliptic functions, e.g. the Weierstrass p-function.

math.CV

CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving

End-to-end planning methods are the de facto standard of the current autonomous driving system, while the robustness of the data-driven approaches suffers due to the notorious long-tail problem (i.e., rare but safety-critical failure cases). In this work, we explore whether recent diffusion-based video generation methods (a.k.a. world models), paired with structured 3D layouts, can enable a fully automated pipeline to self-correct such failure cases. We first introduce an agent to simulate the role of product manager, dubbed PM-Agent, which formulates data requirements to collect data similar to the failure cases. Then, we use a generative model that can simulate both data collection and annotation. However, existing generative models struggle to generate high-fidelity data conditioned on 3D layouts. To address this, we propose DriveSora, which can generate spatiotemporally consistent videos aligned with the 3D annotations requested by PM-Agent. We integrate these components into our self-correcting agentic system, CorrectAD. Importantly, our pipeline is an end-to-end model-agnostic and can be applied to improve any end-to-end planner. Evaluated on both nuScenes and a more challenging in-house dataset across multiple end-to-end planners, CorrectAD corrects 62.5% and 49.8% of failure cases, reducing collision rates by 39% and 27%, respectively.

cs.CV

Exploring Collaboration Breakdowns Between Provider Teams and Patients in Post-Surgery Care

Post-surgery care involves ongoing collaboration between provider teams and patients, which starts from post-surgery hospitalization through home recovery after discharge. While prior HCI research has primarily examined patients' challenges at home, less is known about how provider teams coordinate discharge preparation and care handoffs, and how breakdowns in communication and care pathways may affect patient recovery. To investigate this gap, we conducted semi-structured interviews with 13 healthcare providers and 4 patients in the context of gastrointestinal (GI) surgery. We found coordination boundaries between in- and out-patient teams, coupled with complex organizational structures within teams, impeded the "invisible work" of preparing patients' home care plans and triaging patient information. For patients, these breakdowns resulted in inadequate preparation for home transition and fragmented self-collected data, both of which undermine timely clinical decision-making. Based on these findings, we outline design opportunities to formalize task ownership and handoffs, contextualize co-temporal signals, and align care plans with home resources.

cs.HC

Deep Learning for Clouds and Cloud Shadow Segmentation in Methane Satellite and Airborne Imaging Spectroscopy

Effective cloud and cloud shadow detection is a critical prerequisite for accurate retrieval of concentrations of atmospheric methane (CH4) or other trace gases in hyperspectral remote sensing. This challenge is especially pertinent for MethaneSAT, a satellite mission launched in March 2024, to fill a significant data gap in terms of resolution, precision and swath between coarse-resolution global mappers and fine-scale point-source imagers of methane, and for its airborne companion mission, MethaneAIR. MethaneSAT delivers hyperspectral data at an intermediate spatial resolution (approx. 100 x 400, m), whereas MethaneAIR provides even finer resolution (approx. 25 m), enabling the development of highly detailed maps of concentrations that enable quantification of both the sources and rates of emissions. In this study, we use machine learning methods to address the cloud and cloud shadow detection problem for sensors with these high spatial resolutions. Cloud and cloud shadows in remote sensing data need to be effectively screened out as they bias methane retrievals in remote sensing imagery and impact the quantification of emissions. We deploy and evaluate conventional techniques-including Iterative Logistic Regression (ILR) and Multilayer Perceptron (MLP)-with advanced deep learning architectures, namely U-Net and a Spectral Channel Attention Network (SCAN) method. Our results show that conventional methods struggle with spatial coherence and boundary definition, affecting the detection of clouds and cloud shadows. Deep learning models substantially improve detection quality: U-Net performs best in preserving spatial structure, while SCAN excels at capturing fine boundary details... Our data and code is publicly available at: https://doi.org/10.7910/DVN/IKLZOJ

cs.CV

SegAssess: Panoramic quality mapping for robust and transferable unsupervised segmentation assessment

High-quality image segmentation is fundamental to pixel-level geospatial analysis in remote sensing, necessitating robust segmentation quality assessment (SQA), particularly in unsupervised settings lacking ground truth. Although recent deep learning (DL) based unsupervised SQA methods show potential, they often suffer from coarse evaluation granularity, incomplete assessments, and poor transferability. To overcome these limitations, this paper introduces Panoramic Quality Mapping (PQM) as a new paradigm for comprehensive, pixel-wise SQA, and presents SegAssess, a novel deep learning framework realizing this approach. SegAssess distinctively formulates SQA as a fine-grained, four-class panoramic segmentation task, classifying pixels within a segmentation mask under evaluation into true positive (TP), false positive (FP), true negative (TN), and false negative (FN) categories, thereby generating a complete quality map. Leveraging an enhanced Segment Anything Model (SAM) architecture, SegAssess uniquely employs the input mask as a prompt for effective feature integration via cross-attention. Key innovations include an Edge Guided Compaction (EGC) branch with an Aggregated Semantic Filter (ASF) module to refine predictions near challenging object edges, and an Augmented Mixup Sampling (AMS) training strategy integrating multi-source masks to significantly boost cross-domain robustness and zero-shot transferability. Comprehensive experiments demonstrate that SegAssess achieves state-of-the-art (SOTA) performance and exhibits remarkable zero-shot transferability to unseen masks. The code is available at https://github.com/Yangbn97/SegAssess.

cs.CV

Interfacial reconstruction effects in insulating double perovskite Nd$_2$NiMnO$_6$/SrTiO$_3$ and Nd$_2$NiMnO$_6$/NdGaO$_3$ thin films

Ferromagnetic insulating (FMI) double perovskite oxides (DPOs) $A_2BB'$O$_6$ with near-room-temperature Curie temperatures are promising candidates for ambient-temperature spintronics applications. To realize their potential, epitaxial stabilization of DPO films and understanding the effect of multiple broken symmetries across the film/substrate interface are crucial. This study investigates ultrathin films of the FMI Nd$_2$NiMnO$_6$ (NNMO) grown on SrTiO$_3$ (STO) and NdGaO$_3$ (NGO) substrates. By comparing growth on these substrates, we examine the influence of polarity and structural symmetry mismatches, which are absent in the NGO system. The interface exhibits immeasurable resistance in both cases. Using synchrotron X-ray diffraction, we show that films have three octahedral rotational domains because of the structural symmetry mismatch with the STO substrate. Furthermore, our coherent Bragg rod analysis of specular X-ray diffraction reveals a significant modification of the out-of-plane lattice parameter within a few unit cells at the film/substrate interface and the surface. This arises from polarity compensation and surface symmetry breaking, respectively. These structural alterations influence the Mn orbital symmetry, a dependence that we further confirm through X-ray linear dichroism measurements. Since the ferromagnetism in insulating DPOs is mediated by orbital-dependent superexchange interactions [Phys. Rev. Lett. 100, 186402 (2008)], our study provides a framework for understanding the evolution of magnetism in ultrathin geometry.

cond-mat.mtrl-sci

Balancing Efficiency and Empathy: Healthcare Providers' Perspectives on AI-Supported Workflows for Serious Illness Conversations in the Emergency Department

Serious Illness Conversations (SICs), discussions about values and care preferences for patients with life-threatening illness, rarely occur in Emergency Departments (EDs), despite evidence that early conversations improve care alignment and reduce unnecessary interventions. We interviewed 11 ED providers to identify challenges in SICs and opportunities for technology support, with a focus on AI. Our analysis revealed a four-stage SIC workflow (identification, preparation, conduction, documentation) and barriers at each stage, including fragmented patient information, limited time and space, lack of conversational guidance, and burdensome documentation. Providers expressed interest in AI systems for synthesizing information, supporting real-time conversations, and automating documentation, but emphasized concerns about preserving human connection and clinical autonomy. This tension highlights the need for technologies that enhance efficiency without undermining the interpersonal nature of SICs. We propose design guidelines for ambient and peripheral AI systems to support providers while preserving the essential humanity of these conversations.

cs.HC

Terahertz-field activation of polar skyrons

Unraveling collective modes arising from coupled degrees of freedom is crucial for understanding complex interactions in solids and developing new functionalities. Unique collective behaviors emerge when two degrees of freedom, ordered on distinct length scales, interact. Polar skyrmions, three-dimensional electric polarization textures in ferroelectric superlattices, disrupt the lattice continuity at the nanometer scale with nontrivial topology, leading to previously unexplored collective modes. Here, using terahertz-field excitation and femtosecond x-ray diffraction, we discovered subterahertz collective modes, dubbed 'skyrons', which appear as swirling patterns of atomic displacements functioning as atomic-scale gearsets. Momentum-resolved time-domain measurements of diffuse scattering revealed an avoided crossing in the dispersion relation of skyrons. We further demonstrated that the amplitude and dispersion of skyrons can be controlled by sample temperature and electric-field bias. Atomistic simulations and dynamical phase-field modeling provided microscopic insights into the three-dimensional crystallographic and polarization dynamics. The discovery of skyrons and their coupling with terahertz fields opens avenues for ultrafast control of topological polar structures.

cond-mat.mtrl-sci

Probing the Collision Geometry via Two-Photon Processes in Heavy-Ion Collisions

The initial collision geometry, including the reaction plane, is crucial for interpreting collective phenomena in relativistic heavy-ion collisions, yet it remains experimentally inaccessible through conventional measurements. Recent studies propose utilizing photon-induced processes as a direct probe, leveraging the complete linear polarization of emitted photons whose orientation strongly correlates with the collision geometry. In this work, we employ a QED-based approach to systematically investigate dilepton production via two-photon processes in heavy-ion collisions at RHIC and LHC energies and detector acceptances. Our calculations reveal that dilepton emission exhibits significant sensitivity to the initial collision geometry through both the azimuthal angles of their emission (defined by the relative momentum vector of the two leptons) and the overall momentum orientation of the dilepton pairs. These findings highlight the potential of two-photon-generated dileptons as a novel, polarization-driven probe to quantify the initial collision geometry and reduce uncertainties in characterizing quark-gluon plasma properties.

hep-ph

Influence of the residual magnetic field on the azimuthal distribution of final-state particles in photon-nuclear processes

In relativistic heavy-ion collisions, charged particles are accelerated to nearly the speed of light, and their external electromagnetic fields can be effectively approximated as quasi-real photons. These photons interact with another nucleus via photon-nuclear interactions, producing vector mesons. These vector mesons possess extremely low transverse momentum (pT ~ 0.1 GeV/c), distinguishing them from particles produced via hadronic interactions. STAR and ALICE have observed J/psi, rho0 and other vector mesons with very low pT, which are well described by photoproduction models. This unique characteristic of having extremely low transverse momentum allows them to serve as a novel experimental probe. Recent STAR results show that the equivalent photons in photoproduction processes are fully linearly polarized, affecting the azimuthal distribution of final-state particles like rho0 -> pi+ pi-. Since the polarization links to the initial collision geometry, the rho0 azimuthal modulation can probe nuclear structure. However, the post-collision magnetic field may deflect these particles, distorting the azimuthal distribution and complicating structure measurements. We simulated the distribution of residual magnetic fields over time under different collision conditions using UrQMD for Au+Au collisions at sqrt(sNN)=200 GeV and calculated their effects on the azimuthal modulation ( ) of photoproduced rho0. Our results show that in peripheral collisions, the field significantly alters the for photoproduced rho0 with pT ~ 0.1 GeV/c. This provides key insights for future nuclear structure studies via photoproduction in peripheral collisions.

hep-ph