SearcharxivSearch

arXiv subjects

Miao Liu

Publications and source records attributed to Miao Liu.

At least 19 recordsLinked to original sources

Extremal problems for cancellative and locally thin hypergraphs

We study Tur\'an-type extremal problems for cancellative and locally thin uniform hypergraphs. An $r$-uniform hypergraph is $t$-cancellative if $(\cup_{i=1}^t A_i)\cup B\ne (\cup_{i=1}^t A_i)\cup C$ whenever $A_1,\ldots,A_t,B,C$ are distinct edges. Let $C_t(n,r)$ denote the maximum number of edges in such a hypergraph on $n$ vertices. For all fixed integers $t,k\ge2$, we prove that $C_{2(t-1)}(n,tk)=(1+o(1))\frac{\binom{n}{k}}{\binom{tk-1}{k-1}}$ as $n\to\infty$. In the case $t=2$, this shows that F\"uredi's 2012 upper bound for $C_2(n,2k)$ is asymptotically sharp. The lower bound uses locally sparse induced packings, while the upper bound follows from double counting and a matching argument. More generally, for integers $s\ge t\ge1$, an $r$-uniform hypergraph is locally $(s,t)$-thin if among any $s$ distinct edges, at least $t$ contain a vertex that lies in none of the other $s-1$ edges. This notion includes cancellative hypergraphs as special cases. We establish general upper and lower bounds for the corresponding extremal numbers and determine their polynomial order of growth under suitable divisibility assumptions.

math.CO

Accelerating dynamic simulations of photoexcited materials and their evolution by electron-informed machine learning

Nonadiabatic coupled electron-nuclear dynamics upon electronic excitation underpin the microscopic mechanism and rational modulation of diverse photoinduced functional phenomena in materials, yet their direct first-principles simulations remain computationally demanding. Here we develop a framework for nonadiabatic excited-state machine-learning molecular dynamics (EMLMD) simulations, where the nonequilibrium electronic information upon photoexcitation such as electron temperature is rigorously calibrated from high-precision real-time time-dependent density functional theory (rt-TDDFT) benchmark simulations, enabling accurate reconstruction of excited-state potential energy surfaces (PES). This framework natively incorporates the excited-state electron-phonon couplings and intrinsically captures photoinduced phonon anharmonicity, both of which are missing in standard machine learning molecular dynamics, thus delivering first-principles-level accuracy for excited-state atomic evolutions. Large-scale EMLMD simulations resolve time- and momentum-resolved phonon dynamics in photoexcited materials, directly uncovering the competition between photogenerated coherent phonons and thermal phonons during photoinduced phase transition of bismuth. It also simultaneously resolves elusive atomic-scale microscopic dynamics and global structural rearrangement for selenium photoamorphization. Balancing high accuracy and efficiency, EMLMD offers a versatile paradigm to tackle key challenges in the study of complex excited-state molecular dynamics.

cond-mat.mtrl-sci

Probing strange-sector flavor-changing neutral currents at DUNE-like facilities

We investigate the sensitivity of DUNE-like facilities to flavor-changing neutral-current interactions between down and strange quarks through inclusive deep-inelastic neutrino scattering. In the framework of a general low-energy effective Lagrangian for Dirac and Majorana neutrinos, we evaluate projected sensitivities for CP- and $\tau$-optimized beam configurations. The strongest bounds on the dimensionless Wilson coefficients, defined relative to the Fermi constant $G_F$, reach $\mathcal{O}(10^{-3})\text{--}\mathcal{O}(10^{-2})$ at the near detector, whereas far-detector limits are roughly one order of magnitude weaker. When systematic uncertainties are included, the near-detector sensitivity is dominated by the background-rate uncertainty, while the far-detector sensitivity is limited primarily by the signal-rate uncertainty. At the far detector, the dependence of the projected bounds on the neutrino mass ordering and the Dirac CP-violating phase $\delta_{\mathrm{CP}}$ is largely confined to coefficients involving the electron flavor.

hep-ph

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.

cs.CL

MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models

Object-centric world models forecast future videos by evolving a set of entity slots, but the variables receiving dynamics supervision are often unconstrained visual features. We introduce \method{}, a mask-grounded soft-Hamiltonian world model that makes its position-like state explicitly depend on slot-owned image support. A frozen video-slot encoder produces slots and masks; spatial moments of mask-owned support form a canonical state $Q$, temporal differences form $P$, and a learned energy supplies a soft directional bias to a bounded learned increment. Decoder-relevant appearance and identity are stored separately in a causal visual context. A gated composer and bounded residual then combine this context with the propagated phase state to reconstruct decoder-compatible slots. On OBJ3D, given six observed frames and evaluated over the following 30 frames, \method{} reduces LPIPS by 25.0\% and spatial MSE by 33.7\% relative to the strongest object-centric baseline. On CLEVRER, given six observed frames and evaluated over the following ten frames, the corresponding reductions are 14.5\% and 18.7\%. Horizon-resolved visual and object-state measurements show that the complete model accumulates error more slowly throughout the 30-frame closed-loop rollout. Project page:https://github.com/moshwm-anon/-moshwm-anon.github.io.

cs.LG

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by multimodal assistants,avatars, and embodied agents. While recent audio-video generative models can synthesizehigh-fidelity synchronized content, existing supervision is largely \emph{descriptive}:models are trained to render captions rather than to produce audio-visual responsescaused by external user interactions. We introduce \textbf{InteracVid}, \emph{the firstopen-source large-scale dataset that addresses this missing supervision}, so that everysample couples a preceding audio-visual context and an external stimulus with the realinteractive response that follows. We design a metadata-aware pipeline that extractsinteractive clips from long, noisy livestreams, yielding over \textbf{454K}context-query-response triplets from more than \textbf{59K} livestream videos andspanning conversation-centered, object-centric, procedural, embodied, and screen-basedscenarios. A ten-rater human study confirms that the extracted interactions are causal,natural, and temporally complete for both genuine and reconstructed queries. On aheld-out benchmark of \textbf{100} genuine live-chat queries, fine-tuning on InteracVidimproves both interaction planning and audio-video response generation, and anindependent human evaluation reproduces the system ranking and the conclusions obtainedwith our automatic judge. These results highlight interaction-structured data as acritical foundation for interactive multimodal generation.

cs.CV

Aligning Heterogeneous DFT Datasets: A Graph Neural Network Approach to Cross-Functional Formation Energies

Heterogeneous density functional theory (DFT) calculations, particularly plane-wave implementations, introduce systematic formation energy errors ranging from tens to hundreds of meV/atom, depending on the selection of exchange-correlation functionals, kinetic energy cutoffs, pseudopotentials, and dispersion corrections. As demonstrated by the MatPES dataset, identical structures can exhibit an average energy discrepancy of 107 meV/atom between PBE and r2SCAN calculations. Such method-dependent discrepancies hinder the integration of multi-source DFT data, greatly limiting the scale and quality of datasets for training robust materials AI models. Here, we resolve this fundamental data silo barrier via graph-based transfer learning. Leveraging 380,190 structurally paired PBE-r2SCAN entries from the MatPES database, we train a structure-aware graph neural network to predict cross-functional energy residuals and align inconsistent DFT energy scales. By adopting GPTFF model architecture, the model converts conventional PBE energies to r2SCAN-level accuracy with a mean absolute error of 14.3 meV/atom, compared with 18.2 meV/atom achieved by CHGNet. This versatile approach effectively upgrades massive legacy PBE datasets to high-precision r2SCAN standards. It enables reliable predictions of phase stability, battery voltage profiles, and reaction thermodynamics, while allowing the integration of multi-source DFT data to advance the development of high-performance materials foundation models.

cond-mat.mtrl-sci

MatDiffract: A Material-Informed Automated Analysis Platform for X-ray Powder Diffraction

High-throughput experimentation and self-driving laboratories are drastically accelerating materials discovery, yet automated interpretation of X-ray powder diffraction (XRPD) data remains a critical rate-limiting step. Conventional search-match workflows rely heavily on expert manual intervention, while pure data-driven machine learning approaches suffer from limited generalizability across chemical systems and lack rigorous crystallographic interpretability. Here we present MatDiffract, a material-informed automated analysis platform for high-throughput XRPD characterization. Built on a first-principles density functional theory (DFT)-derived inorganic crystal structure database, Atomly, MatDiffract constructs a perturbation-augmented simulated diffraction database, embeds multi-scale diffraction features into indexable vectors, and integrates hierarchical vector retrieval with full-pattern fitting Rietveld refinement and quantitative phase fitting. Benchmarked on 875 single-phase experimental patterns, the platform achieves 91.3% Top-1 and 97.2% Top-10 identification accuracy after automated refinement. For binary and ternary multiphase mixtures, it delivers 85.0% and 70.0% Top-1 accuracy with mass fraction mean absolute errors as low as 1.2% and 1.8%, respectively. Beyond mere phase labeling, MatDiffract outputs full crystallographic results including refined structural models, fitted profiles, and quantitative compositions within tens of seconds per sample. Its modular vector-based architecture supports seamless incremental expansion to new material systems, providing an end-to-end solution to close the characterization throughput gap for autonomous materials discovery and high-throughput materials development.

cond-mat.mtrl-sci

Graph Neural Network Force Fields (GPTFF-mol) for Organic Molecules from Optimization Trajectories (OpenGEM26)

Density functional theory (DFT) serves as a reliable tool for atomistic molecular simulations, while machine learning potentials have become powerful complements to balance accuracy and efficiency. In this work, we release OpenGEM26 (Open Generated Ensemble of Molecules, 2026), a large-scale dataset comprising 200,000 unique molecules and 4.4 million conformations composed of H, C, N, O, S and Cl with up to ten heavy atoms. All calculations are carried out at the {\omega}B97X-D/Def2-SVP and Def2-TZVP levels with dispersion corrections, and complete structural optimization trajectories and abundant non-equilibrium structures are recorded. Statistical analyses confirm that this dataset covers a broader conformational space than QM9 in terms of energy, bond lengths and bond angles. A graph neural network-based potential GPTFF-mol is trained using the new dataset, achieving an energy mean absolute error of 16 meV/molecule, which is equivalent to 0.82meV/atom, and superior force prediction performance compared with ANI-2x. Validated by butane rotation and keto-enol tautomerization tests, the model accurately describes molecular dynamical behaviors and reaction barriers at distorted geometries. This work provides a high-quality resource and robust ML potential for efficient simulations of sulfur- and chlorine-containing organic molecules.

physics.chem-ph

ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing

Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundational mismatch between the distribution-mapping nature of generative models and the single-point optimization of standard TTT. In this paper, we demonstrate that this mismatch triggers \textit{Prior Collapse}, a degenerate state where the model discards the text conditions and spatial latents, collapsing generations to the source video, or entangling the features of distinct regions. To resolve this, we propose \textbf{ElasticTTT}, a novel framework that preserves the prior generative distribution and rescues generative elasticity. Specifically, we propose \textit{Target Distribution Regularization} to prevent sharp memorization minima, \textit{Contrastive CFG} to guide inference away from source biases, and \textit{Asynchronous Noise Schedule} to preserve unedited regions. Extensive evaluations, supported by theoretical analysis, demonstrate that ElasticTTT successfully preserves the generative prior of the base model, achieving state-of-the-art performance on one-shot video editing.

cs.CV

Upper-shadow comparisons on the slice and the Frankl--Tokushige product conjectures

Let $\mathcal{F}\subseteq\binom{[n]}k$, and let $\partial_{k\to\ell}\mathcal{F}$ be the family of all $\ell$-sets containing a member of $\mathcal{F}$. Writing $\mu_k(\mathcal{F})=|\mathcal{F}|/\binom nk$, $p=k/n$, and $q=\ell/n$, we prove $$ \mu_\ell(\partial_{k\to\ell}\mathcal{F})\geq \begin{cases}\mu_k(\mathcal{F})^{\frac{\log q}{\log p}}, &0\leq \mu_k(\mathcal{F})\leq p^2,\\ q\left[1-\left(1-\frac{\mu_k(\mathcal{F})}{p}\right)^{ \frac{\log(1-q)}{\log(1-p)}} \ \ \right], &p^2\leq \mu_k(\mathcal{F})\leq p,\\ 1-(1-\mu_k(\mathcal{F}))^{\frac{\log(1-q)}{\log(1-p)}}, &p\leq \mu_k(\mathcal{F})\leq1. \end{cases} $$ This yields a dimension-free closed-form lower bound for the finite Kruskal--Katona profile and each branch is asymptotically sharp. As the main application, we settle both the uniform and biased product conjectures of Frankl and Tokushige. If $r\geq2$, $0\leq k_i\leq(r-1)n/r$, and $\mathcal{F}_i\subseteq\binom{[n]}{k_i}$ are $r$-cross-intersecting, then $$ \prod_{i=1}^r\mu_{k_i}(\mathcal{F}_i)\leq\prod_{i=1}^r\frac{k_i}{n}. $$ In biased product setting, we prove that for $r$-cross-intersecting families $\mathcal{A}_i\subseteq 2^{[n]}$ and $0\leq p_i\leq(r-1)/r$, $\prod_{i=1}^r\mu_{p_i}(\mathcal{A}_i)\leq\prod_{i=1}^rp_i.$ We further determine all equality cases in both settings.

math.CO

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics. We introduce MeetingToM, a benchmark for complex social behavior reasoning in naturalistic multi-party meetings. MeetingToM targets meeting-specific phenomena such as \textbf{pseudo-consensus}, where apparent agreement masks private dissent under social pressure. The benchmark is hierarchically organized to evaluate ToM at increasing levels of social granularity, including (i) subject-level mental state prediction, (ii) dyadic-level addressee understanding, and (iii) group-level consensus reasoning. We provide a unified evaluation protocol and conduct systematic analyses of representative MLLMs, revealing persistent limitations in integrating non-verbal cues, inferring hidden attitudes, and distinguishing genuine consensus from pseudo-consensus. Our results highlight key challenges and establish MeetingToM as a testbed for advancing meeting-grounded ToM in multimodal models.

cs.CL

Are Machine Learning Interatomic Potentials Truly Practical? A Benchmark of 23 Mainstream Models

Most MLIP benchmarks reward static accuracy while ignoring inference efficiency and hardware scalability -- driving model bloat with unclear real-world value. We benchmark 23 mainstream open-source MLIPs on a low-cost NVIDIA DGX Spark (128 GB native memory, capped at 80 GB to mimic ordinary lab hardware), using a fixed 192-atom system under a unified ASE-based pipeline. We evaluate three dimensions: predictive accuracy, MD simulation throughput, and atomic scalability. Our results expose a sharp accuracy-efficiency trade-off: large SOTA models deliver only 3-5 meV/atom more accuracy than lightweight ones, but lose orders of magnitude in throughput -- in the worst case, becoming only marginally faster than DFT itself. Lightweight MLIPs, by contrast, sit on the Pareto frontier and run on modest hardware. The lesson is that single-dimensional benchmarks mislead the field, and that future MLIP development should value efficiency and scalability alongside accuracy.

cond-mat.mtrl-sci

PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation

Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personalized image generation, where models must align outputs with a user's implicit visual preferences based on a few historically preferred images and a short prompt. To this end, we introduce PIPBench, the first profile-inclusive benchmark for evaluating personalized image generation. We further propose a novel data construction pipeline that leverages psychological and demographic profiling dimensions for both real-user data collection and scalable agent-based data generation. Using PIPBench, we conduct a thorough evaluation of representative line of methods. Our experiments reveal key limitations in existing methods, suggesting new challenges and opportunities for personalized text-to-image synthesis. Project page: https://wuyuhang05.github.io/PIPBench/

cs.CV

Quench of chiral superconductivity by quantum phase fluctuations in twisted cuprate bilayers

Following theoretical proposals of chiral $d+id'$ superconductivity in twisted cuprate bilayers, experimental signatures of time-reversal symmetry breaking (TRSB) remain highly controversial. Here we demonstrate that quantum phase fluctuations fundamentally reshape the phase diagram of this proposed chiral state. Unlike regular superconducting orders, the chiral $d+id'$ state requires long-range coherence of an interlayer phase degree of freedom and is therefore intrinsically vulnerable to phase fluctuations. Incorporating these fluctuations nearly eliminates the chiral phase over most parts of the phase diagram, restricting it to a narrow twist-angle window and ultra-low temperatures. Phase fluctuations also strongly weaken Josephson phase locking near a twist angle of $45^\circ$. More broadly, our work establishes quantum phase fluctuations as a fundamental constraint on the emergence of TRSB phases in low-dimensional layered quantum materials.

cond-mat.supr-con

Collective phase modes in twisted $d$-wave superconducting bilayers

Twisted cuprate bilayers have been predicted to host high-temperature chiral $d+id'$ superconductivity, originating from higher-order Josephson coupling processes. In such two-dimensional superconducting systems, long-wavelength fluctuations in the phase of the superconducting order parameter constitute gapless collective modes and therefore remain significant even at zero temperature. Here, we perform a theoretical analysis of the low-energy phase fluctuations in twisted $d$-wave superconducting bilayers within a self-consistent harmonic approximation, systematically retaining Josephson coupling to all orders. We demonstrate that higher-order Josephson coupling processes lead to nontrivial modifications of the phase dynamics. The momentum-resolved summand of the relative-phase stiffness is nonzero even in the normal state because interlayer tunneling explicitly breaks the intralayer U(1) symmetry, but its momentum integral vanishes for the continuum dispersion. The relative-phase stiffness is smaller in the $d+id'$ phase than in the $d$-wave phase, while the overall-phase stiffness has the opposite behavior. Furthermore, phase fluctuations strongly soften the Josephson plasma frequency near a twist angle of $45^\circ$ and also substantially reduce the Josephson critical current.

cond-mat.supr-con

The sharp diagonal spectral correlation inequality on the discrete cube

We prove the sharp diagonal spectral correlation conjecture of Friedgut, Kahn, Kalai and Keller, proposed in their Fourier-analytic approach to Chv\'atal's conjecture. For every pair of increasing Boolean functions $f,g:\{0,1\}^n\to\{0,1\}$, $$\mathrm{Cov}(f,g)\ge4\sum_{\varnothing\ne S\subseteq[n]}|S|\hat{f}(S)^2\hat{g}(S)^2.$$ Thus covariance controls the degree-weighted collision of the two nonconstant Fourier spectra, giving a sharp Fourier strengthening of the Harris--Kleitman inequality. The theorem also implies the unweighted diagonal conjecture of Friedgut--Kahn--Kalai--Keller for an increasing family and a maximal intersecting family. The factor $4$ is optimal, and we determine all equality cases. Apart from pairs whose relevant coordinate sets are disjoint, equality occurs only for a common dictatorship and, up to relabelling coordinates and interchanging $f$ and $g$, for the two-coordinate AND-OR pair $(f,g)=(x_i x_j,\,x_i\vee x_j).$ The main novelty is a correlated four-restriction induction and a sharp endpoint convolution inequality. The usual two-restriction induction behind Harris--Kleitman sees only the parallel restricted pairs and loses the mixed Fourier information needed to control the degree-weighted diagonal spectral energy. We instead couple the four codimension-one restricted pairs with correlation $1/2$; this precise correlation extracts the missing degree-weighted energy as a nonnegative square.

math.CO

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding

We introduce EgoSAT, the first comprehensive benchmark for egocentric video reasoning in streaming settings, designed to evaluate the capabilities of modern vision-language models (VLMs). The benchmark targets streaming interaction understanding, where video frames arrive sequentially and models must continuously interpret evolving visual context. EgoSAT unifies several previously distinct tasks within a single streaming framework. In this formulation, queries about completed events correspond to retrospective reasoning, queries about ongoing activities require online understanding, and queries about future actions involve prospective anticipation. This unified setting requires models to reason about the past, present, and future while operating under the constraint that only previously observed frames are available. EgoSAT contains 1,997 unique videos spanning 165 hours of egocentric footage and around 4,800 high-quality question-answer pairs, carefully designed to probe reasoning across varying temporal contexts. Using this benchmark, we evaluate a diverse set of both open-weight and closed-weight VLMs, providing a systematic assessment of their ability for streaming interaction understanding. By distinguishing answerability and conducting diagnostics on confidence of models, we find existing models not only struggle with prospective and retrospective modeling, but also exhibit severe mis-calibration: confidence often fails to track inherent answerability, leading to dangerous "confidently wrong" behaviors. Project page: https://leiyj23.github.io/EgoSAT/

cs.CV