SearcharxivSearch

arXiv subjects

Songtao Li

Publications and source records attributed to Songtao Li.

3 recordsLinked to original sources

Multimodal Alignment and Fusion: A Survey

This survey provides a comprehensive overview of recent advances in multimodal alignment and fusion within the field of machine learning, driven by the increasing availability and diversity of data modalities such as text, images, audio, and video. Unlike previous surveys that often focus on specific modalities or limited fusion strategies, our work presents a structure-centric and method-driven framework that emphasizes generalizable techniques. We systematically categorize and analyze key approaches to alignment and fusion through both structural perspectives -- data-level, feature-level, and output-level fusion -- and methodological paradigms -- including statistical, kernel-based, graphical, generative, contrastive, attention-based, and large language model (LLM)-based methods, drawing insights from an extensive review of over 260 relevant studies. Furthermore, this survey highlights critical challenges such as cross-modal misalignment, computational bottlenecks, data quality issues, and the modality gap, along with recent efforts to address them. Applications ranging from social media analysis and medical imaging to emotion recognition and embodied AI are explored to illustrate the real-world impact of robust multimodal systems. The insights provided aim to guide future research toward optimizing multimodal learning systems for improved scalability, robustness, and generalizability across diverse domains.

cs.CV

Latent Iterative Refinement Flow: A Geometric Constrained Approach for Few-Shot Generation

Diffusion and flow-matching models trained with limited data often tend to memorize the training data instead of generalization, leading to severely reduced diversity. In this paper, we provide a dynamical perspective and identify this ``collapse-to-memorization'' phenomenon as a consequence of the \emph{velocity field collapse}, where the learned field degenerates into isolated point attractors and trap the sampling trajectories. Inspired by this novel view, we introduce \textbf{{\BLUE L}atent {\BLUE I}terative {\BLUE R}efinement {\BLUE F}low ({\BLUE LIRF})}, a geometry-aware framework for from-scratch training of diffusion models in the limited-data regime. By exploiting the intrinsic geometry of a semantically aligned latent space, LIRF progressively densifies the training data manifold via a \emph{generation--correction--augmentation} closed loop, thereby effectively resolving the velocity field collapse. Theoretical guarantee on the convergence of this manifold densification procedure is also provided. Experiments on FFHQ subsets and Low-Shot datasets demonstrate the advantageous performance of LIRF over existing diffusion models for limited-data generation, achieving significantly higher diversity and recall, with comparably good generative performance.

cs.LG

Social norms of fairness with reputation-based role assignment in the dictator game

A vast body of experiments share the view that social norms are major factors for the emergence of fairness in a population of individuals playing the dictator game (DG). Recently, to explore which social norms are conducive to sustaining cooperation has obtained considerable concern. However, thus far few studies have investigated how social norms influence the evolution of fairness by means of indirect reciprocity. In this study, we propose an indirect reciprocal model of the DG and consider that an individual can be assigned as the dictator due to its good reputation. We investigate the `leading eight' norms and all second-order social norms by a two-timescale theoretical analysis. We show that when role assignment is based on reputation, four of the `leading eight' norms, including stern judging and simple standing, lead to a high level of fairness, which increases with the selection intensity. Our work also reveals that not only the correct treatment of making a fair split with good recipients but also distinguishing unjustified unfair split from justified unfair split matters in elevating the level of fairness.

cs.GT