Extract Info

Proud to be supported by Investissement Québec, Scale AI and Confiance IA. Multimodal AI (text & image) applied to product classification in retail. Combining data retrieval, vision-language models, OCR extraction, and embedding-based analysis.

Case study / 04

Project context

Product classification using AI
Recommender systems
Business intelligence for retail

Plumfind contribution

AI & Intelligent Systems
Research & Experimentation
Cloud & Infrastructure

Technology

CLIP, SigLIP, OCR, Embeddings, Vector Search, AWS

Timeline

Discovery, Experimentation, Evaluation, Demo Deployment

An applied research project to explore modern multimodal AI systems for understanding retail products, retrieving visually similar items, and extracting structured information from real-world retail imagery.

Overview / the reason to exist

A useful classification tool starts with a clear question.

What was explored

Multimodal AI approaches for retail product similarity search, model evaluation, OCR extraction, and embedding analysis.

Who it served

Retail technology teams, e-commerce operators, researchers, and AI practitioners.

Why it mattered

Understanding where vision-language models succeed, where they fail, and how complementary systems can improve retail intelligence workflows.

THE CHALLENGE

Retail data is rarely clean. Models must work anyway.

Ecommerce catalogs contain inconsistent photography, varying product descriptions, changing styles, and significant differences between retailers, making them a challenging environment for machine learning systems.


REQUIREMENTS

Reduce uncertainty through experimentation.

Business context

We began with a simple question: how reliably can modern multimodal models understand products across different datasets, stores, and visual conditions?

Approach

We built a collection of practical experiments covering similarity search, CLIP benchmarking, OCR extraction, fine-tuning, store-level analysis, Benchmarked the AI pipeline for new and second-hand products to better understand model performance in realistic conditions.

Experience direction

A research-first approach focused on measurable observations, comparative evaluation, and explainable outcomes rather than production deployment.

Production plan

Lightweight cloud infrastructure supported experimentation, benchmarking, model training, and interactive demonstrations.

discovery

Clarity comes from comparison.

We compared multiple vision-language models, OCR systems, datasets, and store environments to understand how different approaches behave under real-world conditions.

Example research question

“Can modern multimodal models reliably identify visually similar clothing products across different stores?”

  • Evaluated using embedding similarity
  • Compared across multiple CLIP variants
  • Tested against real-world retail imagery
  • Analyzed under domain-shift conditions

What Plumfind designed and evaluated.

A collection of applied AI experiments, benchmark frameworks, retrieval systems, and multimodal analysis workflows.

Content model

Mapped book details, podcast episodes, themes, and metadata into one trusted source layer.


Similarity retrieval

Image embeddings were used to retrieve visually similar products across datasets.

Comparative benchmarking

Multiple CLIP-style models were evaluated against identical datasets to measure consistency and robustness.

OCR intelligence

Structured information was extracted directly from labels, tags, and garment metadata.

Experimental infrastructure

Cloud-based experimentation environments supported training, evaluation, batch processing, and demonstration workflows.

Solution / architecture

Designed as one clear research path.

01

Product image understanding

CLIP-based models generated embeddings capable of representing visual similarity between products.

02

Comparative evaluation

Multiple vision-language models were benchmarked against common product datasets to expose strengths and limitations.

03

OCR extraction

Label information was extracted to provide structured product attributes beyond visual inference.

04

Experimental deployment

Cloud infrastructure enabled reproducible experiments, evaluation pipelines, and interactive demonstrations.


System flow

01

Product imagery

Product images from multiple product datasets and stores.

02

Embedding generation

CLIP-based models generate semantic representations.

03

Similarity & evaluation

Retrieval, comparison, benchmarking, and analysis workflows.

04

Experimental environment

Cloud infrastructure supports testing, training, and demonstrations.

DEVELOPMENT

From research questions to measurable findings.

Experiment design

Data retrieval, OCR extraction, benchmarking, and store-level analysis were structured around practical business and technical questions.

Engineering

Embedding pipelines, similarity search systems, OCR workflows, evaluation frameworks, and fine-tuning experiments formed the technical foundation.

Cloud experimentation

AWS services supported model training, inference endpoints, benchmarking workflows, and demonstration environments.

RESULTS

What became visible through systematic experimentation.

  • Built working product similarity search demonstrations
  • Benchmarked multiple CLIP-style models on product datasets
  • Evaluated OCR systems for attribute extraction
  • Explored store-level embedding analysis
  • Measured domain-shift effects across retailers
  • Tested small-scale CLIP fine-tuning workflows
  • Identified limitations of visual-only product intelligence systems

Key findings

Data quality matters more than model selection.
Cleaning, normalization, and consistency frequently produced larger improvements than changing architectures.
OCR complements vision models.
Structured product information is often more reliable when extracted directly from labels rather than inferred visually.
Retrieval is more robust than classification.
Embedding-based similarity search showed stronger resilience under domain shift than traditional classification tasks.
Retail data is highly variable.
Differences between stores significantly impact model behavior, generalization, and evaluation outcomes.

TECHNOLOGY / SERVICES

Vision-language layer

For image understanding, embedding generation, similarity retrieval, and multimodal evaluation.

OCR layer

For extracting structured product attributes directly from garment labels and tags.

Research layer

For benchmarking models, measuring domain shift, and evaluating experimental outcomes.

Cloud infrastructure

For training, experimentation, deployment, batch processing, and demonstration environments.

Funding context: Supported by Investissement Québec and Scale AI as an applied research and experimentation initiative focused on evaluating multimodal AI capabilities within retail datasets and retail environments.

Related projects

Continue through the work.

Start with the business result.

Bring the question you are trying to answer. We will help you decide what deserves to be built.