Title: Text-guided product image editing based on multimodal feature fusion

Authors: Dianhui Mao; Zhongxin Zheng

Addresses: School of Computer Science and Artificial Intelligence, Beijing Technology and Business University, Beijing, 100048, China; Beijing Key Laboratory of Big Data Technology for Food Safety, Beijing Technology and Business University, Beijing, 100048, China ' School of Computer Science and Artificial Intelligence, Beijing Technology and Business University, Beijing, 100048, China; Beijing Key Laboratory of Big Data Technology for Food Safety, Beijing Technology and Business University, Beijing, 100048, China

Abstract: In order to improve the multimodal consistency and semantic similarity of product image editing results, a text guided product image editing method based on multimodal feature fusion is proposed. Firstly, shape features are extracted through Hu moments, texture characteristics are described with a grey-level co-occurrence matrix, and edge features are detected via the Canny algorithm. Secondly, image features including shape, texture, and edges are integrated with target text information using a dual attention mechanism, thereby achieving multimodal feature fusion. Finally, text guided product image editing is achieved by employing a generative adversarial network model and combining the feature fusion results of target text with existing images. The experimental results demonstrate that a multimodal consistency coefficient of 0.98 and a visual semantic similarity of 0.990 can be achieved by the proposed method.

Keywords: multimodal features; dual attention mechanism; feature fusion; text guided product images; image editing.

DOI: 10.1504/IJCAT.2026.153111

International Journal of Computer Applications in Technology, 2026 Vol.78 No.3, pp.186 - 193

Received: 20 Dec 2024
Accepted: 08 Apr 2025

Published online: 22 Apr 2026 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article