SearcharxivSearch

arXiv subjects

Haohao Liu

Publications and source records attributed to Haohao Liu.

6 recordsLinked to original sources

Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation

We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which only 6.1 billion are active per token. This architecture enables highly efficient scaling (dramatically improving computational efficiency while significantly expanding model capacity) and empowers stronger unified multimodal intelligence across vision, speech, and language, representing a key step toward Artificial General Intelligence (AGI). Compared to its predecessor, the upgraded version exhibits substantial improvements across multimodal understanding and generation. Notably, it achieves strong performance on vision-language understanding benchmarks, with overall scores on par with Gemini 2.5 Pro, and enables seamless switching among multimodal tasks in multi-turn interactions. In speech, it achieves strong performance in contextual and dialect-aware ASR while enabling joint, continuous-generation of speech, sound, and music. In vision, it introduces generative semantic segmentation that achieves competitive standalone performance and enhances spatial control and editing consistency, alongside marked improvements in identity preservation, and high-fidelity in-image text rendering. Together, these capabilities demonstrate that a single unified model can serve as a practical foundation for general-purpose multimodal intelligence.

cs.CV

Variation of Tannaka groups of perverse sheaves in family

Let $k$ be a field of characteristic $0$, let $S$ be a smooth, geometrically connected variety over $k$, with generic point $\eta$, and $f:\mathbb{X}\rightarrow S$ a morphism separated and of finite type. Fix a prime $\ell$. Let $\mathbb{P}$ be an $f$-universally locally acyclic relative perverse $\overline{\mathbb{Q}}_\ell$-sheaf on $\mathbb{X}/S$. We prove that if for some (equivalently, every) geometric point $\bar \eta$ over $\eta$ the restriction $\mathbb{P}|_{\mathbb{X}_{\bar \eta}}$ is simple as a perverse $\overline{\mathbb{Q}}_\ell$-sheaf on $\mathbb{X}_{\bar \eta}$, then there is a non-empty open subscheme $U\subset S$ such that, for every geometric point $\bar s$ on $U$, the restriction $\mathbb{P}|_{\mathbb{X}_{\bar s}}$ is simple as a perverse $\overline{\mathbb{Q}}_\ell$-sheaf on $\mathbb{X}_{\bar s}$. When $f:\mathbb{X}\rightarrow S$ is an abelian scheme, we give applications of this result to the variation with $s\in S$ of the Tannaka group of $\mathbb{P}|_{\mathbb{X}_{\bar s}}$.

math.AG

Normality of monodromy group in generic convolution group

On an abelian variety $A$, sheaf convolution gives a Tannakian formalism for perverse sheaves. Let $X$ be an irreducible algebraic variety with generic point $\eta$. Let $K$ be a family of perverse sheaves (more precisely, a relative perverse sheaf) on the constant abelian scheme $p_X:A\times X\to X$. We show that for uncountably many character sheaves $L_{\chi}$ on $A$, the monodromy groups of $R^0p_{X*}(K\otimes p_A^*L_{\chi})$ are normal in the Tannakian group $G(K|_{A_{\eta}})$ of the perverse sheaf $K|_{A_{\eta}}\in\mathrm{Perv}(A_{\eta})$. This result is inspired from and could be compared to two other normality results: In the same setting, the Tannakian group $G(K|_{A_{\bar{\eta}}})$ is normal in $G(K|_{A_{\eta}})$ (due to Lawrence-Sawin). For a polarizable variation of Hodge structures, outside a meager locus, the connected monodromy group is normal in the derived Mumford-Tate group (due to Andr\'e).

math.AG

Generic vanishing theorem for Fujiki class C

A Nakano-type generic vanishing result is extended from compact Kähler manifolds to manifolds in Fujiki class $\mathcal{C}$, so that smooth proper complex algebraic varieties are covered.

math.AG

Maxwell: a hardware and software highly integrated compute-storage system

The compute-storage framework is responsible for data storage and processing, and acts as the digital chassis of all upper-level businesses. The performance of the framework affects the business's processing throughput, latency, jitter, and etc., and also determines the theoretical performance upper bound that the business can achieve. In financial applications, the compute-storage framework must have high reliability and high throughput, but with low latency as well as low jitter characteristics. For some scenarios such as hot-spot account update, the performance of the compute-storage framework even surfaces to become a server performance bottleneck of the whole business system. In this paper, we study the hot-spot account issue faced by Alipay and present our exciting solution to this problem by developing a new compute-storage system, called Maxwell. Maxwell is a distributed compute-storage system with integrated hardware and software optimizations. Maxwell does not rely on any specific hardware (e.g. GPUs or FPGAs). Instead, it takes deep advantage of computer components' characteristics, such as disk, network, operating system and CPU, and aims to emit the ultimate performance of both hardware and software. In comparison with the existing hot-spot account updating solutions deployed online, Maxwell achieves three orders of magnitude performance improvement for end-to-end evaluation. Meanwhile, Maxwell also demonstrates remarkable performance gains in other related businesses of Ant Group.

cs.DC