Summary

Genesis 01 tests whether vision models recognize designer and season from pixels alone.

Key takeaways

Download PDF
Cite this
@techreport{peace2025genesis01,
  title = {Testing AI Vision's Understanding of High Fashion Nuances},
  author = {Peace, Kalan},
  institution = {Caeliai},
  year = {2025},
  url = {https://research.caeliai.com/research/genesis-01}
}
Fig. 1

SigLIP kept collections together; CLIP and DINOv2 did not.

SigLIP63.5%
CLIP48.8%
DINOv248.6%

Collection purity by model.

Source: Genesis 01 paper table (research/genesis-01.tex).

Uncertainty detection gap

SigLIP+0.079
CLIP-0.015
DINOv2-0.070

Positive values indicate lower confidence on impostors than on true matches. Shared scale: −0.100 to +0.100.

Source: genesis-01.tex.

Needle precision

SigLIP100%
CLIP90.0%
DINOv271.4%

Exact-match precision as published in the complete results table.

Source: genesis-01.tex.

Back to research

Genesis 01 · Research paper

Testing AI vision’s understanding of high fashion nuances.

Kalan Peace · Caeliai · Published in 2025

Also published as “A Vision Benchmark for Fashion AI”.

The first study compares how CLIP, SigLIP, and DINOv2 handle fashion images, focusing on uncertainty, impostor detection, and collection cohesion.

Using 12,147 Rick Owens runway images across 23 years, the study runs 3.66 million image comparisons. In this evaluation, SigLIP is the only model that becomes more cautious when the image is ambiguous.

What the study shows.

The study is less about raw similarity scores and more about behavior: whether a model can recognize a design family, stay coherent across a collection, and mark uncertainty when it should.

Scope

Genesis 01 compares three vision backbones across three tasks: impostor detection, collection cohesion, and exact look matching. The goal is not generic image classification. It is how the models handle uncertain fashion images.

3 Models compared
3 Benchmark tasks
2002-2025 Temporal range
Key result
In this evaluation, SigLIP is the only model that becomes more cautious when the image is an impostor.

CLIP and DINOv2 remain more confident on the ambiguous cases. SigLIP is more expensive, but it shows the strongest uncertainty behavior in this evaluation.

63.5% Collection purity
9.6x Higher processing cost

The benchmark is built from three concrete scenes.

Each scene maps to a real product problem: false similarity, weak collection understanding, or failure to retrieve the precise object the user means.

01 / Impostor detection

From pixels alone, the model has to recognize that visually adjacent looks are still the wrong designer or wrong season and should trigger uncertainty.

Genesis 01 query image
Query
Genesis 01 impostor example one
Impostor A
Genesis 01 impostor example two
Impostor B
02 / Collection cohesion

One runway look should retrieve its family, not a group of vaguely similar silhouettes from different years. This separates collection matching from general visual similarity.

Genesis 01 family resemblance query
Anchor look
Genesis 01 family resemblance match one
True family
Genesis 01 family resemblance contamination example
Contamination
03 / Exact match

Retrieval also has to stay precise. When a user points at a specific look, the system should find that exact target instead of drifting toward a general aesthetic neighborhood.

Genesis 01 exact match target
Target
Genesis 01 found match
Returned match
Genesis 01 haystack comparison image
Hard negative
Operational read
A vision model needs to show when it is unsure.

If the model cannot mark uncertainty, its output may favor a visually similar but incorrect reference. That is a model-behavior issue, not only a score.

Why it matters

Recommendation systems may use visual similarity to describe products. If a model cannot separate designer identity from visual similarity, its output may become an overconfident approximation.