\documentclass[10.5pt, a4paper]{article} \usepackage[T1]{fontenc} \usepackage[top=0.75in, bottom=0.75in, left=0.85in, right=0.85in]{geometry} \usepackage{xcolor} \usepackage{hyperref} \usepackage{booktabs} \usepackage{array} \usepackage{enumitem} \usepackage{titlesec} \usepackage{graphicx} \usepackage{lmodern} \usepackage{amssymb} \usepackage{listings} \usepackage{mdframed} \usepackage{multirow} \usepackage{colortbl} % Colors \definecolor{sage}{HTML}{BAD9B5} \definecolor{charcoal}{HTML}{2A2A2C} \definecolor{offwhite}{HTML}{EBEBEB} \definecolor{codebg}{HTML}{F5F5F5} % Hyperlinks \hypersetup{colorlinks=true, linkcolor=charcoal, urlcolor=charcoal, pdftitle={Testing AI Vision's Understanding of High Fashion Nuances}} % Section formatting \titleformat{\section}{\normalsize\bfseries\color{charcoal}}{}{0em}{\MakeUppercase} \titlespacing*{\section}{0pt}{10pt}{4pt} \titleformat{\subsection}{\normalsize\bfseries}{}{0em}{} \titlespacing*{\subsection}{0pt}{6pt}{2pt} \setlength{\parindent}{0pt} \setlength{\parskip}{4pt} \renewcommand{\arraystretch}{1.2} % Code listing style \lstset{ basicstyle=\ttfamily\fontsize{8}{10}\selectfont, backgroundcolor=\color{codebg}, frame=single, framerule=0.4pt, rulecolor=\color{charcoal}, breaklines=true, breakatwhitespace=true, showstringspaces=false, commentstyle=\color{charcoal}, keywordstyle=\bfseries, language=Python, } % Key finding box \newmdenv[ linecolor=charcoal, linewidth=1pt, leftline=true, rightline=true, topline=true, bottomline=true, backgroundcolor=offwhite, innerleftmargin=16pt, innerrightmargin=16pt, innertopmargin=12pt, innerbottommargin=12pt, ]{findingbox} % Sage accent box \newmdenv[ linecolor=sage, linewidth=2pt, leftline=true, rightline=false, topline=false, bottomline=false, innerleftmargin=12pt, innerrightmargin=0pt, innertopmargin=6pt, innerbottommargin=6pt, ]{sagebox} % ───────────────────────────────────────────── \begin{document} % HEADER \noindent \begin{minipage}[t]{0.6\textwidth} \includegraphics[height=20pt]{../public/caeliai-Offical-Logo.png}\hspace{6pt}{\large\textbf{Caeliai}}\\[2pt] {\small\color{charcoal} Research / Genesis-01} \end{minipage} \hfill \begin{minipage}[t]{0.35\textwidth} \raggedleft {\small\color{charcoal} Published: 2025\\ \href{mailto:contact@caeliai.com}{contact@caeliai.com} } \end{minipage} \vspace{6pt} \noindent\rule{\textwidth}{2pt} \vspace{10pt} % TITLE {\fontsize{20}{24}\selectfont\textbf{Testing AI Vision's Understanding\\[4pt]of High Fashion Nuances}} \vspace{6pt} {\small\color{charcoal}\textit{The first comprehensive evaluation of vision models for fashion intelligence}} \vspace{6pt} \noindent{\color{sage}\rule{\linewidth}{3pt}} \vspace{10pt} % ABSTRACT \section{Abstract} We present the first comprehensive benchmark evaluating vision models for fashion intelligence, testing \texttt{CLIP}, \texttt{SigLIP}, and \texttt{DINOv2} on \texttt{12,147} Rick Owens runway images across \texttt{23} years of collections. Through \texttt{3.66 million} image comparisons across three evaluation frameworks --- impostor detection, collection cohesion, and exact matching --- we demonstrate that \texttt{SigLIP} achieves superior fashion intelligence with \texttt{+0.079} positive uncertainty detection gap and \texttt{63.5\%} collection purity, making it the only model suitable for production fashion AI systems despite \texttt{9.6x} processing cost. \vspace{8pt} \begin{findingbox} \begin{center} {\large\textbf{Key Finding}}\\[8pt] \textbf{SigLIP} is the only model with proper uncertainty detection, achieving \textbf{+0.079} positive noise gap while \texttt{CLIP} (\textcolor{charcoal}{$-0.015$}) and \texttt{DINOv2} (\textcolor{charcoal}{$-0.070$}) show dangerous overconfidence. \end{center} \end{findingbox} \vspace{10pt} % METHODOLOGY \section{1. Methodology} \subsection{1.1 Model Selection} We evaluated three state-of-the-art vision models representing different architectural approaches to multimodal understanding: \begin{itemize}[leftmargin=*, itemsep=3pt] \item \textbf{SigLIP} (\texttt{google/siglip-so400m-patch14-384}): Google's enhanced CLIP variant with sigmoid loss, \texttt{1152}-dimensional embeddings, \texttt{400M} parameters. \item \textbf{CLIP} (\texttt{openai/clip-vit-base-patch32}): OpenAI's original contrastive model, \texttt{512}-dimensional embeddings. \item \textbf{DINOv2} (\texttt{facebookresearch/dinov2\_vitb14}): Meta's self-supervised vision model, \texttt{768}-dimensional embeddings. \end{itemize} \subsection{1.2 Dataset Construction} Our dataset comprises \texttt{12,147} Rick Owens runway images spanning \texttt{23} years (\texttt{2002--2025}) across multiple collections: \begin{sagebox} \textbf{Fall Collections:} 2002, 2003, 2004, 2008, 2013, 2015, 2016, 2021, 2024\\ \textbf{Spring Collections:} 2004, 2006, 2007, 2010, 2012, 2013, 2015, 2016, 2017, 2019, 2020, 2023, 2025\\ \textbf{Coverage:} Menswear, Ready-to-Wear, Beauty campaigns \end{sagebox} \subsection{1.3 Technical Infrastructure} Processing performed on \texttt{M4 Mac} with \texttt{Apple Silicon GPU}. Embeddings stored in \texttt{Pinecone} vector database for similarity search. All models implemented using \texttt{PyTorch} with \texttt{transformers} library. \vspace{8pt} % THREE CORE TESTS \section{2. Three Core Intelligence Tests} \renewcommand{\arraystretch}{1.3} \begin{tabular}{>{\bfseries}p{0.22\textwidth} p{0.35\textwidth} p{0.35\textwidth}} Test & Task & What good looks like \\ \midrule Impostor Detection & From pixels only, identify if candidates belong to the same designer and season/look as the query & Proper \emph{uncertainty} on impostors --- not overconfident matching \\[6pt] Family Recognition & Given one look from a season, find other looks from the same designer and season & High collection purity --- low cross-season contamination \\[6pt] Needle Search & Find the exact same look in a database of 12,147 images & Rank 1 match = exact designer, season, and look number \\ \end{tabular} \renewcommand{\arraystretch}{1.2} \vspace{10pt} % PERFORMANCE METRICS \section{3. Performance Metrics \& Trade-offs} \subsection{3.1 Uncertainty Detection} Positive gap = model is properly less confident on impostors than true matches. Only SigLIP passes. \vspace{6pt} \begin{tabular}{>{\bfseries}p{0.15\textwidth} p{0.22\textwidth} p{0.22\textwidth} p{0.3\textwidth}} Model & Uncertainty Gap & Verdict & Interpretation \\ \midrule SigLIP & \textbf{+0.079} & \textcolor{black}{PASS} & Proper uncertainty on impostors \\ CLIP & $-0.015$ & \textcolor{charcoal}{FAIL} & Overconfident on noise \\ DINOv2 & $-0.070$ & \textcolor{charcoal}{FAIL} & Dangerously overconfident \\ \end{tabular} \subsection{3.2 Collection Purity} \begin{tabular}{>{\bfseries}p{0.15\textwidth} p{0.22\textwidth} p{0.5\textwidth}} Model & Purity & Note \\ \midrule SigLIP & \textbf{63.5\%} & Maintains season cohesion \\ CLIP & 48.8\% & Cross-season contamination \\ DINOv2 & 48.6\% & Near-random collection grouping \\ \end{tabular} \subsection{3.3 Processing Speed} Designer Gauntlet Test (50 comparisons): \begin{tabular}{>{\bfseries}p{0.15\textwidth} p{0.15\textwidth} p{0.55\textwidth}} Model & Speed & Trade-off \\ \midrule CLIP & 1.9 min & Fast but unreliable \\ DINOv2 & 2.7 min & Fast but unreliable \\ SigLIP & 18.7 min & 9.6x cost, only production-ready option \\ \end{tabular} \vspace{10pt} % TECHNICAL IMPLEMENTATION \section{4. Technical Implementation} \subsection{4.1 SigLIP Embedding Generation} \begin{lstlisting} from transformers import AutoModel, AutoProcessor import torch # Load SigLIP model siglip_model = AutoModel.from_pretrained( "google/siglip-so400m-patch14-384", torch_dtype=torch.float16 ).to("mps") # Apple Silicon GPU # Generate embedding image = Image.open("rick-owens-fall-2011-menswear-39.jpg") inputs = siglip_processor(images=image, return_tensors="pt") with torch.no_grad(): image_features = siglip_model.get_image_features(**inputs) embedding = image_features / image_features.norm(dim=-1, keepdim=True) \end{lstlisting} \subsection{4.2 Vector Search Implementation} \begin{lstlisting} async def identify_image(self, image_embedding: List[float]): """Find most similar runway image using cosine similarity.""" # Query Pinecone vector database query_result = self.img_index.query( vector=image_embedding, top_k=1, include_metadata=True, filter={"file_type": "look"} # Only runway looks ) # Calculate confidence similarity_score = match.score if similarity_score >= 0.95: confidence = 99.0 # Near-perfect match elif similarity_score >= 0.90: confidence = 95.0 # Strong match \end{lstlisting} \vspace{8pt} % COMPLETE RESULTS \section{5. Complete Results Summary} \begin{tabular}{>{\bfseries}p{0.12\textwidth} p{0.22\textwidth} p{0.18\textwidth} p{0.18\textwidth} p{0.2\textwidth}} \toprule Model & Uncertainty Detection & Collection Purity & Needle Precision & Processing Speed \\ \midrule \rowcolor{offwhite} SigLIP & +0.079 gap (PASS) & 63.5\% & 100\% accuracy & 18.7 min (10x cost) \\ CLIP & $-0.015$ gap (FAIL) & 48.8\% & 90.0\% accuracy & 3.1 s \\ DINOv2 & $-0.070$ gap (FAIL) & 48.6\% & 71.4\% accuracy & 4.2 s \\ \bottomrule \end{tabular} \vspace{10pt} % CONCLUSION \section{Conclusion} SigLIP is the only vision model suitable for production fashion AI. Its \texttt{+0.079} uncertainty gap and \texttt{63.5\%} collection purity demonstrate genuine semantic understanding of designer aesthetic, not just visual similarity matching. CLIP and DINOv2's negative uncertainty gaps --- meaning they are \emph{more} confident on impostors than true matches --- make them categorically unsuitable for high-precision fashion retrieval. The \texttt{9.6x} processing cost of SigLIP is a real constraint but acceptable at production scale for high-value fashion intelligence applications. The benchmark establishes SigLIP at \texttt{google/siglip-so400m-patch14-384} as the foundational model for any serious fashion AI system requiring reliable similarity search across large runway archives. \vspace{12pt} \noindent\rule{\textwidth}{2pt} \vspace{6pt} \noindent{\small\color{charcoal}Caeliai Research $\cdot$ \href{https://caeliai.com/research}{caeliai.com/research} $\cdot$ \href{mailto:contact@caeliai.com}{contact@caeliai.com}} \end{document}