Publication Date

7-23-2026

Document Type

Article

Publication Title

Machine Learning with Applications

Volume

25

DOI

10.1016/j.mlwa.2026.100961

Abstract

Large language models and vision-language models enable new forms of multimodal representation learning, yet recognizing abstract fashion styles remains challenging. Most existing models focus on category-level clothing recognition and rely primarily on visual features, which often fail to capture higher-level stylistic semantics such as casual, chic, or sporty. This work introduces StyleFusion, a multimodal learning framework for fashion style recognition that integrates visual representations with semantic cues derived from LLM-generated descriptions. Image embeddings extracted using CLIP are combined with style-aware textual representations generated through caption prompting and encoded with vanilla BERT. A cross-attention fusion mechanism aligns visual and textual representations to capture complementary information relevant to style perception. The proposed architecture is evaluated on the FashionStyle14 dataset containing 14 fashion style categories. Three multimodal fusion strategies: feature concatenation, gated fusion, and cross-attention fusion are compared under a shared encoder backbone to isolate the contribution of the fusion mechanism. Experimental results show that cross-attention fusion achieves slightly higher mean performance than the image-only frozen CLIP baseline, the text-only model, and simpler multimodal fusion strategies, although the observed improvement over the frozen CLIP baseline is not statistically significant. The proposed model achieves performance comparable to the image-only fine-tuned CLIP baseline. The proposed model achieves an accuracy of 0.83 and an F1 score of 0.83, maintaining stable performance across multiple random seeds. These findings suggest that LLM-generated descriptions provide complementary cues for abstract fashion style recognition, which is more prominent when the visual encoder is frozen, while the benefit becomes limited when the visual encoder is already fine-tuned for the target task. This study contributes to an empirical analysis of cross-attention multimodal fusion for fashion style recognition and clarifies when the LLM-generated descriptions provide useful complementary information for abstract fashion style recognition.

Keywords

Attentionbased fusion, Deep learning, Fashion style classification, Image-text representation, Multimodal Learning, Vision-language models

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution-Noncommercial 4.0 License

Department

Applied Data Science

Share

COinS