SearcharxivSearch

arXiv subjects

Ying Cao

Publications and source records attributed to Ying Cao.

At least 19 recordsLinked to original sources

Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates

In this paper, we address the problem of graphic design template creation, which generates a background image and a layout of foreground elements over the background to form a harmonious composition from an input text. Prior work on graphic design generation mostly adopts a sequential paradigm, where design elements are generated sequentially. We argue that such a sequential scheme falls short of faithfully capturing the dependency between the background and layout (and thus the joint image-layout distribution), which limits the quality of generated design templates. To overcome this limitation, we propose a model, InterIL, which jointly generates the two modalities, background image and layout, in a single generative process. The novel design of our joint model connects the backbones of pretrained image and layout diffusion models with a learnable communication module to explicitly model bidirectional image-layout interaction. During training, the image and layout backbones are frozen to maintain and leverage the vast pretrained single-modality prior knowledge, while only the communication module is updated, so that the model can focus on learning image-layout interaction and thereby better capture the joint image-layout distribution for improved composition harmony. Our model has no design-specific inductive bias, which allows it to better preserve the original characteristics of realistic designs. We further introduce a test-time guidance strategy to enable users to impose their specific preferences on generated results. Our experiments show that, compared with prior approaches, our model can generate significantly better results in terms of image, layout and image-layout harmonization, producing outputs closer to real samples. We also demonstrate the flexibility of our model in enforcing user preferences at inference without retraining.

cs.CV

Inverse reinforcement learning for indefinite mean-field social optimization with multiplicative noise

This paper studies the inverse reinforcement learning (RL) problem for linear-quadratic mean-field (MF) social optimization. The considered system features multiplicative noise and indefinite cost weights, which violate standard convexity assumptions and pose analytical challenges. The goal is to recover unknown social cost weights from expert demonstrations and reproduce the optimal control policies. This requires solving coupled stochastic algebraic Riccati equations and Lyapunov equations with unknown system dynamics. To this end, we first propose a model-based inverse RL algorithm with two sequential loops that separately handle individual and MF dynamics, and we prove its convergence and closed-loop stabilizability. Moreover, we characterize the non-uniqueness of the recovered cost weights. To eliminate reliance on system dynamics, we develop a model-free inverse RL algorithm using integral RL and least-squares identification, which requires only measured trajectory data satisfying mild rank conditions. Finally, numerical simulations validate the effectiveness of the proposed approaches.

math.OC

Stabilizer Design for Policy Iteration in Stochastic Linear Quadratic Control: A Spectrum-Assignment Approach

Policy iteration (PI) is an important reinforcement learning tool for solving optimal control problems which includes an initialization stage, i.e., the search for an initial stabilizing controller. However, the initialization stage typically relies on complete model information, thereby imposing substantial constraints on the initialization of model-free PI. For stochastic systems with multiplicative noise dependent on state and control, the stability is not ensured by Hurwitz conditions as in the deterministic case, but rather by a Lyapunov-type inequality that incorporates both drift and diffusion terms. Therefore, the corresponding model-free PI initialization problem is more challenging. To this end, a novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control. With the help of the Lyapunov-type operator's spectrum, the original system is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtaining a stabilizing control gain. Furthermore, by leveraging system data and adjusting the cumulative factor, we design a model-free algorithm that does not rely on an initial stabilizing policy and can achieve optimal control. Finally, simulation results are provided to validate the effectiveness of the proposed methods.

math.OC

Partition-selected flow polynomials and associated arrangements

We introduce a partition-selection method to generalize the flow, chromatic, and Tutte polynomials of a graph by restricting the standard edge subgraph expansions to subgraphs given by prescribed connected vertex partitions. We establish similar deletion-contraction formulas and specialization relations for these polynomials, recovering all classical polynomial invariants when the selection is the set of all partitions. Next we study a relation between Jaeger et al.'s nonhomogeneous flows and a special class of partition-selected flow polynomials (called affine flow polynomials). Specifically, we give a geometric realization of nowhere-zero nonhomogeneous flows by restricting the edge-coordinate arrangement to affine flow spaces. The resulting characteristic polynomials coincide with Kochol's admissible assigning polynomials and with affine flow polynomials, which enumerate nowhere-zero nonhomogeneous flows over finite fields. To see the key role of the partition-selection framework, we further introduce boundary arrangements determined by the bond structure of a graph. Using the intersection posets of boundary arrangements, we obtain the classification of all restricted arrangements mentioned above, the comparison of unsigned coefficients of affine flow polynomials, and the decomposition formulas for affine flow polynomials.

math.CO

Ehrhart quasi-polynomials of rational polytopes by real dilations

This paper is to study the Ehrhart function $L(P,t)$ of a rational $n$-polytope $P$, defined as the number of lattice points of dilated polytopes $tP$ with real numbers $t\geq 0$. It turns out that $L(P,t)$ is a quasi-polynomial of real variable $t$ in the sense that \[ L(P,t)=\sum_{k=0}^{n} c_k(P,t)t^k, \quad t\geq 0, \] where $c_k(P,t)$ are periodic piecewise polynomials of degree $n-k$ if ${\rm aff}\,P$ contains the origin, and are periodic functions vanishing almost everywhere otherwise. When $P$ is a rational simplex $\sigma$, the coefficient functions $c_k(\sigma,t)$ are given explicitly in terms of vertex information of the simplex $\sigma$. Moreover, the reciprocity law still holds.

math.CO

Elder-Sim: A Psychometrically Validated Platform for Personality-Stable Elderly Digital Twins

Background: LLMs enable patient-facing conversational agents, creating a pathway toward digital twins that capture older adults' lived experiences and behavioral responses across time. A central barrier is personality drift -- inconsistent trait expression across repeated interactions -- which undermines reliability of generated trajectories and intervention-response simulation in geriatric care. Objective: To develop ELDER-SIM, a multi-role elderly-care conversational platform for building personality-stable digital twin agents, and to propose a psychometric validation framework for quantifying personality consistency in LLM-based agents. Methods: ELDER-SIM was implemented via n8n workflow orchestration with local LLM inference (Ollama/vLLM), integrating (1) Big Five (OCEAN) trait specifications, (2) a Cognitive Conceptualization Diagram (CCD) grounded in Beck's CBT framework, and (3) a MySQL-based long-term memory module. Ablation studies across four conditions -- Baseline, +Memory, +CCD, and +LoRA (fine-tuned on 19,717 instruction pairs from CHARLS) -- were evaluated via Cronbach's $\alpha$, ICC, and role discrimination accuracy. Results: Reliability was acceptable to excellent across conditions (Cronbach's $\alpha$: 0.70--0.94; ICC: 0.85--0.96). Role discrimination improved from 83.3% (Baseline) to 88.9% (+Memory), 94.4% (+CCD), and 97.2% (+LoRA). CCD produced the largest consistency gain (mean $\alpha$ 0.702$\to$0.892), while LoRA achieved the highest overall consistency ($\alpha$ 0.940; ICC 0.958). Conclusions: ELDER-SIM provides a psychometrically validated approach for constructing personality-consistent elderly digital twin agents. Structured cognitive modeling and domain adaptation reduce personality drift, supporting reliable longitudinal simulation for elderly mental health care and reproducible in silico evaluation before clinical deployment.

cs.HC

Characteristic quasi-polynomials of truncated arrangements

Given an (affine) integral arrangement $\mathcal{A}$ in $\mathbb{R}^n$, the reduction of $\mathcal{A}$ modulo an arbitrary positive integer $q$ naturally yields an arrangement $\mathcal{A}_q$ in $\mathbb{Z}_q^n$. Our primary objective is to study the combinatorial aspects of the restriction $\mathcal{A}^{(B,\bm b)}$ to the solution space of $B\bm x=\bm b$, and its reduction $\mathcal{A}_q^{(B,\bm b)}$ modulo $q$. This work generalizes the earlier results of Kamiya, Takemura and Terao, as well as Chen and Wang. The purpose of this paper is threefold as follows. Firstly, we derive an explicit counting formula for the cardinality of the complement $M\big(\mathcal{A}_q^{(B,\bm b)}\big)$ of $\mathcal{A}_q^{(B,\bm b)}$; and prove that for all positive integers $q>q_0$, this cardinality coincides with a quasi-polynomial $\chi^{\text{quasi}}\big(\mathcal{A}^{(B,\bm b)},q\big)$ in $q$ with a period $\rho_C$. Secondly, we weaken Chen and Wang's original hypothesis $a \mid b$ to a strictly more general condition $\gcd(a,\rho_C)\mid \gcd(b,\rho_C)$, and introduce the concept of combinatorial equivalence for positive integers. Within this framework, we establish three unified comparison relations: between the unsigned coefficients of $\chi^{\text{quasi}}\big(\mathcal{A}^{(B,\bm b)},a\big)$ and $\chi^{\text{quasi}}\big(\mathcal{A}^{(B,\bm b)},b\big)$; between the unsigned coefficients of distinct constituents of $\chi^{\text{quasi}}\big(\mathcal{A}^{(B,\bm b)},q\big)$; and between the cardinalities of $M\big(\mathcal{A}_q^{(B,\bm b)}\big)$ and $M\big(\mathcal{A}_{pq}^{(B,\bm b)}\big)$. Thirdly, using our method, we revisit the enumerative aspects of group colorings and nowhere-zero nonhomogeneous form flows from the early work of Forge, Zaslavsky and Kochol.

math.CO

CD-DPE: Dual-Prompt Expert Network Based on Convolutional Dictionary Feature Decoupling for Multi-Contrast MRI Super-Resolution

Multi-contrast magnetic resonance imaging (MRI) super-resolution intends to reconstruct high-resolution (HR) images from low-resolution (LR) scans by leveraging structural information present in HR reference images acquired with different contrasts. This technique enhances anatomical detail and soft tissue differentiation, which is vital for early diagnosis and clinical decision-making. However, inherent contrasts disparities between modalities pose fundamental challenges in effectively utilizing reference image textures to guide target image reconstruction, often resulting in suboptimal feature integration. To address this issue, we propose a dual-prompt expert network based on a convolutional dictionary feature decoupling (CD-DPE) strategy for multi-contrast MRI super-resolution. Specifically, we introduce an iterative convolutional dictionary feature decoupling module (CD-FDM) to separate features into cross-contrast and intra-contrast components, thereby reducing redundancy and interference. To fully integrate these features, a novel dual-prompt feature fusion expert module (DP-FFEM) is proposed. This module uses a frequency prompt to guide the selection of relevant reference features for incorporation into the target image, while an adaptive routing prompt determines the optimal method for fusing reference and target features to enhance reconstruction quality. Extensive experiments on public multi-contrast MRI datasets demonstrate that CD-DPE outperforms state-of-the-art methods in reconstructing fine details. Additionally, experiments on unseen datasets demonstrated that CD-DPE exhibits strong generalization capabilities.

cs.CV

Using the Schmidt Decomposition to Determine Quantum Entanglement

Quantum information theory is a rapidly growing area of math and physics that combines two independent theories, quantum mechanics and information theory. Quantum entanglement is a concept that was first proposed in the EPR paradox. In quantum mechanics, particles can be in superposition, meaning they are in multiple different states at once. It is not until the particle is measured that it is forced into a single state. However, it is possible that particles can be tied to other particles, meaning that the measurement of one particle will determine the measurement of the other particle. Entanglement is at the very core of quantum information theory. It is one of the core pieces that allows for the massive increase in computing power. For this paper, we decided to focus on demonstrating the mathematical method (the Schmidt decomposition) for determining if a system is entangled, and a demonstration of quantum entanglement's use (quantum teleportation) as well as a quick look at how to extend the uses of the Schmidt decomposition.

quant-ph

Breaking the Likelihood Trap: Consistent Generative Recommendation with Graph-structured Model

Reranking, as the final stage of recommender systems, plays a crucial role in determining the final exposure, directly influencing user experience. Recently, generative reranking has gained increasing attention for formulating reranking as a holistic sequence generation task, implicitly modeling complex dependencies among items. However, most existing methods suffer from the likelihood trap, where high-likelihood sequences are often repetitive and perceived as low-quality by humans, thereby limiting user engagement. In this work, we propose Consistent Graph-structured Generative Recommendation (CONGRATS). We first introduce a novel Graph-structured Model, which enables the generation of more diverse sequences by exploring multiple paths. This design not only expands the decoding space to promote diversity, but also improves prediction accuracy by explicitly modeling item dependencies from graph transitions. Furthermore, we design a Consistent Differentiable Training method that incorporates an evaluator, allowing the model to learn directly from user preferences. Extensive offline experiments validate the superior performance of CONGRATS over state-of-the-art reranking methods. Moreover, CONGRATS has been evaluated on a large-scale video-sharing app, Kuaishou, with over 300 million daily active users, demonstrating that our approach significantly improves both recommendation quality and diversity, validating our effectiveness in practical industrial platforms.

cs.IR

Stability and Generalization of Adversarial Diffusion Training

Algorithmic stability is an established tool for analyzing generalization. While adversarial training enhances model robustness, it often suffers from robust overfitting and an enlarged generalization gap. Although recent work has established the convergence of adversarial training in decentralized networks, its generalization properties remain unexplored. This work presents a stability-based generalization analysis of adversarial training under the diffusion strategy for convex losses. We derive a bound showing that the generalization error grows with both the adversarial perturbation strength and the number of training steps, a finding consistent with single-agent case but novel for decentralized settings. Numerical experiments on logistic regression validate these theoretical predictions.

cs.LG

On the Escaping Efficiency of Distributed Adversarial Training Algorithms

Adversarial training has been widely studied in recent years due to its role in improving model robustness against adversarial attacks. This paper focuses on comparing different distributed adversarial training algorithms--including centralized and decentralized strategies--within multi-agent learning environments. Previous studies have highlighted the importance of model flatness in determining robustness. To this end, we develop a general theoretical framework to study the escaping efficiency of these algorithms from local minima, which is closely related to the flatness of the resulting models. We show that when the perturbation bound is sufficiently small (i.e., when the attack strength is relatively mild) and a large batch size is used, decentralized adversarial training algorithms--including consensus and diffusion--are guaranteed to escape faster from local minima than the centralized strategy, thereby favoring flatter minima. However, as the perturbation bound increases, this trend may no longer hold. In the simulation results, we illustrate our theoretical findings and systematically compare the performance of models obtained through decentralized and centralized adversarial training algorithms. The results highlight the potential of decentralized strategies to enhance the robustness of models in distributed settings.

cs.LG

Emergent Explicit Regulation in College Students Collaborative Scientific Inquiry Learning, Framework and A Case Study

Small group activities have been widely adopted in college level science courses. As students participate in these activities, it is important to consider how group members collectively regulate their activity and complete group task. Regulation in a group often involves adaptive responsivity from group members when they notice and deal with a challenge. The theoretical framework of socially shared regulation emphasizes group members collaboratively regulating within the group but does not focus on portraying how the shared regulation is developed in the moment. Currently, the field lacks a framework characterizing the momentary development of a regulatory action in a group. Our study addresses this gap. In our video data, incoming college students were enrolled in a summer program designed to promote students metacognitive skills to be incorporated in their study of science. We have observed various moments in which the students spontaneously made a move to regulate the hands-on, inquiry activity in completing their tasks and achieving group goals. We developed a framework called Emergent Explicit Regulation to characterize those moments. The EER framework captures students in the moment regulatory moves to respond to a challenge, articulating how those moves emerge, in what ways they are explicit and regulatory. In this paper, we first introduce the EER framework and situate the EER framework in the context of collaborative scientific inquiry learning. We then present a case study where we applied the EER framework to identify typical EER instances in one small group when the students completed the task of building a model to represent the climate of the Earth atmosphere. They worked collaboratively, faced and handled various challenges, completed the group task, and demonstrated multiple EERs in different psychological areas and in the inquiry practices designed in the activity.

physics.ed-ph

VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation

We present VISTAR, a user-centric, multi-dimensional benchmark for text-to-image (T2I) evaluation that addresses the limitations of existing metrics. VISTAR introduces a two-tier hybrid paradigm: it employs deterministic, scriptable metrics for physically quantifiable attributes (e.g., text rendering, lighting) and a novel Hierarchical Weighted P/N Questioning (HWPQ) scheme that uses constrained vision-language models to assess abstract semantics (e.g., style fusion, cultural fidelity). Grounded in a Delphi study with 120 experts, we defined seven user roles and nine evaluation angles to construct the benchmark, which comprises 2,845 prompts validated by over 15,000 human pairwise comparisons. Our metrics achieve high human alignment (>75%), with the HWPQ scheme reaching 85.9% accuracy on abstract semantics, significantly outperforming VQA baselines. Comprehensive evaluation of state-of-the-art models reveals no universal champion, as role-weighted scores reorder rankings and provide actionable guidance for domain-specific deployment. All resources are publicly released to foster reproducible T2I assessment.

cs.CV

A No-Reference Medical Image Quality Assessment Method Based on Automated Distortion Recognition Technology: Application to Preprocessing in MRI-guided Radiotherapy

Objective:To develop a no-reference image quality assessment method using automated distortion recognition to boost MRI-guided radiotherapy precision.Methods:We analyzed 106,000 MR images from 10 patients with liver metastasis,captured with the Elekta Unity MR-LINAC.Our No-Reference Quality Assessment Model includes:1)image preprocessing to enhance visibility of key diagnostic features;2)feature extraction and directional analysis using MSCN coefficients across four directions to capture textural attributes and gradients,vital for identifying image features and potential distortions;3)integrative Quality Index(QI)calculation,which integrates features via AGGD parameter estimation and K-means clustering.The QI,based on a weighted MAD computation of directional scores,provides a comprehensive image quality measure,robust against outliers.LOO-CV assessed model generalizability and performance.Tumor tracking algorithm performance was compared with and without preprocessing to verify tracking accuracy enhancements.Results:Preprocessing significantly improved image quality,with the QI showing substantial positive changes and surpassing other metrics.After normalization,the QI's average value was 79.6 times higher than CNR,indicating improved image definition and contrast.It also showed higher sensitivity in detail recognition with average values 6.5 times and 1.7 times higher than Tenengrad gradient and entropy.The tumor tracking algorithm confirmed significant tracking accuracy improvements with preprocessed images,validating preprocessing effectiveness.Conclusions:This study introduces a novel no-reference image quality evaluation method based on automated distortion recognition,offering a new quality control tool for MRIgRT tumor tracking.It enhances clinical application accuracy and facilitates medical image quality assessment standardization, with significant clinical and research value.

eess.IV

A Novel Automatic Real-time Motion Tracking Method in MRI-guided Radiotherapy Using Enhanced Tracking-Learning-Detection Framework with Automatic Segmentation

Background and Purpose: Accurate motion tracking in MRI-guided Radiotherapy (MRIgRT) is essential for effective treatment delivery. This study aimed to enhance motion tracking precision in MRIgRT through an automatic real-time markerless tracking method using an enhanced Tracking-Learning-Detection (ETLD) framework with automatic segmentation. Materials and Methods: We developed a novel MRIgRT motion tracking and segmentation method by integrating the ETLD framework with an improved Chan-Vese model (ICV), named ETLD+ICV. The ETLD framework was upgraded for real-time cine MRI, including advanced image preprocessing, no-reference image quality assessment, an enhanced median-flow tracker, and a refined detector with dynamic search region adjustments. ICV was used for precise target volume coverage, refining the segmented region frame by frame using tracking results, with key parameters optimized. The method was tested on 3.5D MRI scans from 10 patients with liver metastases. Results: Evaluation of 106,000 frames across 77 treatment fractions showed sub-millimeter tracking errors of less than 0.8mm, with over 99% precision and 98% recall for all subjects in the Beam Eye View(BEV)/Beam Path View(BPV) orientation. The ETLD+ICV method achieved a dice global score of more than 82% for all subjects, demonstrating the method's extensibility and precise target volume coverage. Conclusion: This study successfully developed an automatic real-time markerless motion tracking method for MRIgRT that significantly outperforms current methods. The novel method not only delivers exceptional precision in tracking and segmentation but also shows enhanced adaptability to clinical demands, making it an indispensable asset in improving the efficacy of radiotherapy treatments.

eess.IV

Competitive Analysis of Online Path Selection: Impacts of Path Length, Topology, and System-Level Costs

Consider a communication network to which a sequence of self-interested users come and send requests for data transmission between nodes. This work studies the question of how to guide the path selection choices made by those online-arriving users and maximize the social welfare. Competitive analysis is the main technical tool. Specifically, the impacts of path length bounds and topology on the competitive ratio of the designed algorithm are analyzed theoretically and explored experimentally. We observe intricate and interesting relationships between the empirical performance and the studied network parameters, which shed some light on how to design the network. We also investigate the influence of system-level costs on the optimal algorithm design.

cs.DS

On the Trade-off between Flatness and Optimization in Distributed Learning

This paper proposes a theoretical framework to evaluate and compare the performance of stochastic gradient algorithms for distributed learning in relation to their behavior around local minima in nonconvex environments. Previous works have noticed that convergence toward flat local minima tend to enhance the generalization ability of learning algorithms. This work discovers three interesting results. First, it shows that decentralized learning strategies are able to escape faster away from local minima and favor convergence toward flatter minima relative to the centralized solution. Second, in decentralized methods, the consensus strategy has a worse excess-risk performance than diffusion, giving it a better chance of escaping from local minima and favoring flatter minima. Third, and importantly, the ultimate classification accuracy is not solely dependent on the flatness of the local minimum but also on how well a learning algorithm can approach that minimum. In other words, the classification accuracy is a function of both flatness and optimization performance. In this regard, since diffusion has a lower excess-risk than consensus, when both algorithms are trained starting from random initial points, diffusion enhances the classification accuracy. The paper examines the interplay between the two measures of flatness and optimization error closely. One important conclusion is that decentralized strategies deliver in general enhanced classification accuracy because they strike a more favorable balance between flatness and optimization performance compared to the centralized solution.

cs.LG