Publication Date

Spring 2026

Degree Type

Thesis

Degree Name

Master of Science (MS)

Department

Applied Data Science

Advisor

Mohammad Masum; Guannan Liu; Saptarshi Sengupta

Abstract

Accurate garment attribute recognition is essential for fashion recommendation systems, trend analysis, and large-scale retrieval applications. However, most current approaches or methods are based on rough visual features and fail to consider about fine-grained semantic attributes that determine clothing style and compatibility. This research investigates visual-to-textual garment labeling using CLIP and extends its capabilities through fine-tuning on curated subset of the Polyvore dataset, which contains diverse images of clothing items. We further extend the idea by introducing a new pipeline incorporates linear probe over CLIP’s embedding structure to capture five attributes: category, subcategory, color, material, and pattern, which are the basic elements of clothing. The process involves freezing both image and text encoders attaching lightweight linear heads for classification. This design preserves CLIP’s rich semantic representations while enabling efficient fashion specific learning. While the accuracy performance reach 93% for category and 72% for subcategory, and high semantic similarity scores of 87%, 87%, and 85% for color, material, and pattern. Another contribution is incorporating CLIP-based semantic similarity into the evaluation pipeline, which provides a more nuanced and interpretable assessment of attribute recognition capturing partial correctness and closely aligned with semantics rather than perfect matches. Together, these contributions advance fine-grained fashion understanding and provide a practical and explainable framework for fashion recommendation systems.

Available for download on Saturday, July 29, 2028

Share

COinS