LEXA: LLM Accuracy Evaluation Framework
Proud to be supported by Investissement Québec and Confiance IA. An AI evaluation initiative for assessing the reliability, accuracy, and performance of large language models for commercial applications.
EVALUATION FRAMEWORK / APPLIED AI

LLM EVALUATION / E-COMMERCE DOMAIN
Case study / 07
Project context
Applied AI & E-commerce
Plumfind contribution
AI Model Evals
Metrics & Tools
Analytics & Data
Technology
LLMs, RAG, OCR, evaluation methods and metrics
PROJECT STATUS
Active research / ongoing evaluation

BUSINESS CHALLENGE
Generative AI needs to earn your trust before deployment.
Organizations adopting LLM-powered recommendation and automation systems need a reliable way to measure factual accuracy, detect hallucinations, and evaluate performance in real-world deployments.
LLM RELIABILITY / DEPLOYMENT GAP
DISCOVERY & RESEARCH
01 / Evaluation landscape
We reviewed LLM evaluation, RAG assessment, hallucination detection, and information-extraction methods to define a practical evaluation framework.
02 / Domain dataset
A Shopify-based clothing-catalogue dataset brought the evaluation into a real e-commerce context, with structured product descriptions and metadata.
03 / Baseline evaluation
We tested multiple accuracy metrics for feature extraction, summarization, hallucination, and completeness.
04 / Multimodal enrichment
OCR extracted product information from images and merged it with catalog metadata to test whether multimodal context improves model accuracy.
Research REQUIREMENTS / operational field study
A practical framework for real-world LLMs.
LEXA connects evaluation criteria with domain-specific data, allowing teams to assess answer relevance, hallucination, summarisation accuracy, context use, and RAG performance together.

SOLUTION
From evaluation criteria to a dependable framework.
01
Evaluation dimensions
A clear model of answer relevance, hallucination, summarisation accuracy, context utilisation, and RAG performance made the assessment actionable.
02
Domain-grounded dataset
A retail dataset based on public Shopify product listings ensured the work reflected commercial complexity rather than a generic academic benchmark.
03
Multimodal enrichment
OCR and product-information enrichment tested whether additional context could strengthen model reliability and answer quality.
Architecture / system flow
The system is designed for trustworthy decisions.
Each stage links source data, LLM responses, evaluation criteria, and evidence so teams can understand where performance is strong and where it needs improvement.

Implementation
From model assessment to a dependable evaluation system.
01
Research & benchmark design
A structured review and evaluation plan translated academic methods into an commercially relevant testing framework.
02
LLM & RAG evaluation
Baseline models, RAG behaviour, and multimodal enrichment were evaluated against a domain-specific dataset and practical quality measures.
03
Insights & next steps
Results identify where additional data, better context, or model refinement can improve reliable deployment.
BUSINESS IMPACT
Trustworthy AI teams can act on.
- Defined practical measures for relevance, hallucination, and summarisation quality
- Built a domain-specific e-commerce benchmark dataset
- Evaluated baseline LLM and RAG performance against real-world questions
- Assessed OCR enrichment and context utilisation efficiency
LESSON LEARNED
Trustworthy AI needs evaluation methods that reflect its data, use cases, and risks—not generic benchmarks alone.

TECHNOLOGIES USED
Evaluation framework
AI
For measuring model quality across answer relevance, hallucination, summarisation, context use, and RAG performance.
Domain dataset
ANALYTICS
For grounding model evaluation in real product data, structured descriptions, and commercial context.
Multimodal enrichment
CLOUD
For testing whether OCR and additional information sources improve the quality and reliability of model answers.
Related projects
Continue through the work
START A CONVERSATION
Start with the business result.
Bring the question you are trying to answer. We will help you decide what deserves to be built.

