LEXA: LLM Accuracy Evaluation Framework

Proud to be supported by Investissement Québec and Confiance IA. An AI evaluation initiative for assessing the reliability, accuracy, and performance of large language models for commercial applications.

EVALUATION FRAMEWORK / APPLIED AI

LLM EVALUATION / E-COMMERCE DOMAIN

Case study / 07

Project context

Applied AI & E-commerce

Plumfind contribution

AI Model Evals
Metrics & Tools
Analytics & Data

Technology

LLMs, RAG, OCR, evaluation methods and metrics

PROJECT STATUS

Active research / ongoing evaluation

BUSINESS CHALLENGE

Generative AI needs to earn your trust before deployment.

Organizations adopting LLM-powered recommendation and automation systems need a reliable way to measure factual accuracy, detect hallucinations, and evaluate performance in real-world deployments.


LLM RELIABILITY / DEPLOYMENT GAP

DISCOVERY & RESEARCH

01 / Evaluation landscape

We reviewed LLM evaluation, RAG assessment, hallucination detection, and information-extraction methods to define a practical evaluation framework.

02 / Domain dataset

A Shopify-based clothing-catalogue dataset brought the evaluation into a real e-commerce context, with structured product descriptions and metadata.

03 / Baseline evaluation

We tested multiple accuracy metrics for feature extraction, summarization, hallucination, and completeness.

04 / Multimodal enrichment

OCR extracted product information from images and merged it with catalog metadata to test whether multimodal context improves model accuracy.

Research REQUIREMENTS / operational field study

A practical framework for real-world LLMs.

LEXA connects evaluation criteria with domain-specific data, allowing teams to assess answer relevance, hallucination, summarisation accuracy, context use, and RAG performance together.

SOLUTION

From evaluation criteria to a dependable framework.

01

Evaluation dimensions

A clear model of answer relevance, hallucination, summarisation accuracy, context utilisation, and RAG performance made the assessment actionable.

02

Domain-grounded dataset

A retail dataset based on public Shopify product listings ensured the work reflected commercial complexity rather than a generic academic benchmark.

03

Multimodal enrichment

OCR and product-information enrichment tested whether additional context could strengthen model reliability and answer quality.

Architecture / system flow

The system is designed for trustworthy decisions.

Each stage links source data, LLM responses, evaluation criteria, and evidence so teams can understand where performance is strong and where it needs improvement.

Implementation

From model assessment to a dependable evaluation system.

01

Research & benchmark design

A structured review and evaluation plan translated academic methods into an commercially relevant testing framework.

02

LLM & RAG evaluation

Baseline models, RAG behaviour, and multimodal enrichment were evaluated against a domain-specific dataset and practical quality measures.

03

Insights & next steps

Results identify where additional data, better context, or model refinement can improve reliable deployment.

BUSINESS IMPACT

Trustworthy AI teams can act on.

  • Defined practical measures for relevance, hallucination, and summarisation quality
  • Built a domain-specific e-commerce benchmark dataset
  • Evaluated baseline LLM and RAG performance against real-world questions
  • Assessed OCR enrichment and context utilisation efficiency

LESSON LEARNED

Trustworthy AI needs evaluation methods that reflect its data, use cases, and risks—not generic benchmarks alone.

TECHNOLOGIES USED

Evaluation framework

AI

For measuring model quality across answer relevance, hallucination, summarisation, context use, and RAG performance.

Domain dataset

ANALYTICS

For grounding model evaluation in real product data, structured descriptions, and commercial context.

Multimodal enrichment

CLOUD

For testing whether OCR and additional information sources improve the quality and reliability of model answers.

Related projects

Continue through the work

START A CONVERSATION

Start with the business result.

Bring the question you are trying to answer. We will help you decide what deserves to be built.