Extract Info
Proud to be supported by Investissement Québec, Scale AI and Confiance IA. Multimodal AI (text & image) applied to product classification in retail. Combining data retrieval, vision-language models, OCR extraction, and embedding-based analysis.

Case study / 04
Project context
Product classification using AI
Recommender systems
Business intelligence for retail
Plumfind contribution
AI & Intelligent Systems
Research & Experimentation
Cloud & Infrastructure
Technology
CLIP, SigLIP, OCR, Embeddings, Vector Search, AWS
Timeline
Discovery, Experimentation, Evaluation, Demo Deployment

An applied research project to explore modern multimodal AI systems for understanding retail products, retrieving visually similar items, and extracting structured information from real-world retail imagery.
Overview / the reason to exist
A useful classification tool starts with a clear question.
What was explored
Multimodal AI approaches for retail product similarity search, model evaluation, OCR extraction, and embedding analysis.
Who it served
Retail technology teams, e-commerce operators, researchers, and AI practitioners.
Why it mattered
Understanding where vision-language models succeed, where they fail, and how complementary systems can improve retail intelligence workflows.
THE CHALLENGE
Retail data is rarely clean. Models must work anyway.
Ecommerce catalogs contain inconsistent photography, varying product descriptions, changing styles, and significant differences between retailers, making them a challenging environment for machine learning systems.
REQUIREMENTS
Reduce uncertainty through experimentation.
Business context
We began with a simple question: how reliably can modern multimodal models understand products across different datasets, stores, and visual conditions?
Approach
We built a collection of practical experiments covering similarity search, CLIP benchmarking, OCR extraction, fine-tuning, store-level analysis, Benchmarked the AI pipeline for new and second-hand products to better understand model performance in realistic conditions.
Experience direction
A research-first approach focused on measurable observations, comparative evaluation, and explainable outcomes rather than production deployment.
Production plan
Lightweight cloud infrastructure supported experimentation, benchmarking, model training, and interactive demonstrations.
discovery

Clarity comes from comparison.
We compared multiple vision-language models, OCR systems, datasets, and store environments to understand how different approaches behave under real-world conditions.
Example research question
“Can modern multimodal models reliably identify visually similar clothing products across different stores?”
- Evaluated using embedding similarity
- Compared across multiple CLIP variants
- Tested against real-world retail imagery
- Analyzed under domain-shift conditions
What Plumfind designed and evaluated.
A collection of applied AI experiments, benchmark frameworks, retrieval systems, and multimodal analysis workflows.
Content model
Mapped book details, podcast episodes, themes, and metadata into one trusted source layer.
Similarity retrieval
Image embeddings were used to retrieve visually similar products across datasets.
Comparative benchmarking
Multiple CLIP-style models were evaluated against identical datasets to measure consistency and robustness.
OCR intelligence
Structured information was extracted directly from labels, tags, and garment metadata.
Experimental infrastructure
Cloud-based experimentation environments supported training, evaluation, batch processing, and demonstration workflows.
Solution / architecture
Designed as one clear research path.
01
Product image understanding
CLIP-based models generated embeddings capable of representing visual similarity between products.
02
Comparative evaluation
Multiple vision-language models were benchmarked against common product datasets to expose strengths and limitations.
03
OCR extraction
Label information was extracted to provide structured product attributes beyond visual inference.
04
Experimental deployment
Cloud infrastructure enabled reproducible experiments, evaluation pipelines, and interactive demonstrations.
System flow

01
Product imagery
Product images from multiple product datasets and stores.
02
Embedding generation
CLIP-based models generate semantic representations.
03
Similarity & evaluation
Retrieval, comparison, benchmarking, and analysis workflows.
04
Experimental environment
Cloud infrastructure supports testing, training, and demonstrations.
DEVELOPMENT
From research questions to measurable findings.
Experiment design
Data retrieval, OCR extraction, benchmarking, and store-level analysis were structured around practical business and technical questions.
Engineering
Embedding pipelines, similarity search systems, OCR workflows, evaluation frameworks, and fine-tuning experiments formed the technical foundation.
Cloud experimentation
AWS services supported model training, inference endpoints, benchmarking workflows, and demonstration environments.
RESULTS
What became visible through systematic experimentation.
- Built working product similarity search demonstrations
- Benchmarked multiple CLIP-style models on product datasets
- Evaluated OCR systems for attribute extraction
- Explored store-level embedding analysis
- Measured domain-shift effects across retailers
- Tested small-scale CLIP fine-tuning workflows
- Identified limitations of visual-only product intelligence systems
Key findings
Data quality matters more than model selection.
Cleaning, normalization, and consistency frequently produced larger improvements than changing architectures.
OCR complements vision models.
Structured product information is often more reliable when extracted directly from labels rather than inferred visually.
Retrieval is more robust than classification.
Embedding-based similarity search showed stronger resilience under domain shift than traditional classification tasks.
Retail data is highly variable.
Differences between stores significantly impact model behavior, generalization, and evaluation outcomes.
TECHNOLOGY / SERVICES
Vision-language layer
For image understanding, embedding generation, similarity retrieval, and multimodal evaluation.
OCR layer
For extracting structured product attributes directly from garment labels and tags.
Research layer
For benchmarking models, measuring domain shift, and evaluating experimental outcomes.
Cloud infrastructure
For training, experimentation, deployment, batch processing, and demonstration environments.
Funding context: Supported by Investissement Québec and Scale AI as an applied research and experimentation initiative focused on evaluating multimodal AI capabilities within retail datasets and retail environments.
Related projects
Continue through the work.
Start with the business result.
Bring the question you are trying to answer. We will help you decide what deserves to be built.

