SearcharxivSearch

arXiv subjects

Zhipeng Jin

Publications and source records attributed to Zhipeng Jin.

6 recordsLinked to original sources

UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis

Text-to-image generation has greatly advanced content creation, yet accurately rendering visual text remains a key challenge due to blurred glyphs, semantic drift, and limited style control. Existing methods often rely on pre-rendered glyph images as conditions, but these struggle to retain original font styles and color cues, necessitating complex multi-branch designs that increase model overhead and reduce flexibility. To address these issues, we propose a segmentation-guided framework that uses pixel-level visual text masks -- rich in glyph shape, color, and spatial detail -- as unified conditional inputs. Our method introduces two core components: (1) a fine-tuned bilingual segmentation model for precise text mask extraction, and (2) a streamlined diffusion model augmented with adaptive glyph conditioning and a region-specific loss to preserve textual fidelity in both content and style. Our approach achieves state-of-the-art performance on the AnyText benchmark, significantly surpassing prior methods in both Chinese and English settings. To enable more rigorous evaluation, we also introduce two new benchmarks: GlyphMM-benchmark for testing layout and glyph consistency in complex typesetting, and MiniText-benchmark for assessing generation quality in small-scale text regions. Experimental results show that our model outperforms existing methods by a large margin in both scenarios, particularly excelling at small text rendering and complex layout preservation, validating its strong generalization and deployment readiness.

cs.CV

A cyclical route linking fundamental mechanism and AI algorithm: An example from tuning Poisson's ratio in amorphous networks

"AI for science" is widely recognized as a future trend in the development of scientific research. Currently, although machine learning algorithms have played a crucial role in scientific research with numerous successful cases, relatively few instances exist where AI assists researchers in uncovering the underlying physical mechanisms behind a certain phenomenon and subsequently using that mechanism to improve machine learning algorithms' efficiency. This article uses the investigation into the relationship between extreme Poisson's ratio values and the structure of amorphous networks as a case study to illustrate how machine learning methods can assist in revealing underlying physical mechanisms. Upon recognizing that the Poisson's ratio relies on the low-frequency vibrational modes of dynamical matrix, we can then employ a convolutional neural network, trained on the dynamical matrix instead of traditional image recognition, to predict the Poisson's ratio of amorphous networks with a much higher efficiency. Through this example, we aim to showcase the role that artificial intelligence can play in revealing fundamental physical mechanisms, which subsequently improves the machine learning algorithms significantly.

cond-mat.soft

Enhancing Dynamic Image Advertising with Vision-Language Pre-training

In the multimedia era, image is an effective medium in search advertising. Dynamic Image Advertising (DIA), a system that matches queries with ad images and generates multimodal ads, is introduced to improve user experience and ad revenue. The core of DIA is a query-image matching module performing ad image retrieval and relevance modeling. Current query-image matching suffers from limited and inconsistent data, and insufficient cross-modal interaction. Also, the separate optimization of retrieval and relevance models affects overall performance. To address this issue, we propose a vision-language framework consisting of two parts. First, we train a base model on large-scale image-text pairs to learn general multimodal representation. Then, we fine-tune the base model on advertising business data, unifying relevance modeling and retrieval through multi-objective learning. Our framework has been implemented in Baidu search advertising system "Phoneix Nest". Online evaluation shows that it improves cost per mille (CPM) and click-through rate (CTR) by 1.04% and 1.865%.

cs.IR

Boost CTR Prediction for New Advertisements via Modeling Visual Content

Existing advertisements click-through rate (CTR) prediction models are mainly dependent on behavior ID features, which are learned based on the historical user-ad interactions. Nevertheless, behavior ID features relying on historical user behaviors are not feasible to describe new ads without previous interactions with users. To overcome the limitations of behavior ID features in modeling new ads, we exploit the visual content in ads to boost the performance of CTR prediction models. Specifically, we map each ad into a set of visual IDs based on its visual content. These visual IDs are further used for generating the visual embedding for enhancing CTR prediction models. We formulate the learning of visual IDs into a supervised quantization problem. Due to a lack of class labels for commercial images in advertisements, we exploit image textual descriptions as the supervision to optimize the image extractor for generating effective visual IDs. Meanwhile, since the hard quantization is non-differentiable, we soften the quantization operation to make it support the end-to-end network training. After mapping each image into visual IDs, we learn the embedding for each visual ID based on the historical user-ad interactions accumulated in the past. Since the visual ID embedding depends only on the visual content, it generalizes well to new ads. Meanwhile, the visual ID embedding complements the ad behavior ID embedding. Thus, it can considerably boost the performance of the CTR prediction models previously relying on behavior ID features for both new ads and ads that have accumulated rich user behaviors. After incorporating the visual ID embedding in the CTR prediction model of Baidu online advertising, the average CTR of ads improves by 1.46%, and the total charge increases by 1.10%.

cs.IR

Revealing the three-component structure of water with principal component analysis (PCA) on X-ray spectrum

Combining the principal component analysis (PCA) of X-ray spectrum with MD simulations, we experimentally reveal the existence of three basic components in water. These components exhibit distinct structures, densities, and temperature dependencies. Among the three, two major components correspond to the low-density liquid (LDL) and the high-density liquid (HDL) predicted by the two-component model, and the third component exhibits a unique 5-hydrogen-bond configuration with an ultra-high local density. As the temperature increases, the LDL component decreases and the HDL component increases, while the third component varies non-monotonically with a peak around 20 $^{\circ}$C to 30 $^{\circ}$C. The 3D structure of the third component is further illustrated as the uniform distribution of five hydrogen-bonded neighbors on a spherical surface. Our study reveals experimental evidence for water's unique three-component structure, which provides a fundamental basis for understanding water's special properties and anomalies.

cond-mat.soft

Achieving adjustable elasticity with non-affine to affine transition

For various engineering and industrial applications it is desirable to realize mechanical systems with broadly adjustable elasticity to respond flexibly to the external environment. Here we discover a topology-correlated transition between affine and non-affine regimes in elasticity in both two- and three-dimensional packing-derived networks. Based on this transition, we numerically design and experimentally realize multifunctional systems with adjustable elasticity. Within one system, we achieve solid-like affine response, liquid-like non-affine response and a continuous tunability in between. Moreover, the system also exhibits a broadly tunable Poisson's ratio from positive to negative values, which is of practical interest for energy absorption and for fracture-resistant materials. Our study reveals a fundamental connection between elasticity and network topology, and demonstrates its practical potential for designing mechanical systems and metamaterials.

cond-mat.mtrl-sci