help1
shopify page speedcore web vitalslighthousecruxmethodology

Lab Tests vs Real Users: A Shopify Page Speed Tools Primer

Six shopify page speed numbers for one store, no two alike. Each tool measures a different thing on purpose - here is what each is for.

help1 Team
Lab Tests vs Real Users: A Shopify Page Speed Tools Primer

You pulled six numbers for the same store on the same morning. PageSpeed Insights mobile: 41. PSI desktop: 87. GTmetrix: a "B". Pingdom: 92. Lighthouse in your own DevTools: 58. Shopify admin: a blank panel where a single score used to be. The store did not change between those numbers. None of them are wrong, and none of them measure the same thing.

This post is the methodology primer. It is not the post about merchants getting scammed with screenshots (that one is over here) and it is not the post about catching a cloaking script in theme.liquid (detective story here). This is the one that walks you through what each shopify page speed tool is actually measuring and how to read it on purpose.

Lab tests versus field tests, the only distinction that matters

Every Shopify performance tool falls into one of two buckets, and the buckets answer different questions.

Lab tests run a synthetic browser on a controlled network. They load your page once, in a known device profile, on a known connection. PageSpeed Insights, GTmetrix, Pingdom, and Chrome DevTools' built-in Lighthouse are all lab tests. They answer the question "how does this URL load right now, in this exact simulated environment?" They are diagnostic instruments.

Field tests collect samples from real users on real devices. The Chrome User Experience Report (CrUX) aggregates real-user metrics across the rolling 28 days when those users had your store open in their actual browser. The Shopify Web Performance Dashboard, Google Search Console's Core Web Vitals report, and the public CrUX API all surface the same field data in different views. They answer the question "what did my customers actually experience over the last month?" That is the only question Google's search ranking signal listens to.

Lab and field scores will not match, because they are measuring different things. A lab score is a verdict on one moment in a simulator. A field score is a 28-day p75 across a population of real visitors. Asking "which shopify page speed score is right" is like asking whether your scale or your doctor is right when one says you weigh 170 and the other says you are 5'10". Different question.

What PageSpeed Insights is actually doing

PSI loads your URL inside a headless Chrome instance pretending to be a Moto G4 on a throttled 4G connection. The throttling is a CPU and bandwidth simulation applied after the page loads rather than real network throttling, designed to approximate what a mid-range Android phone on a contended cell network would see.

The output is a 0-to-100 composite score built from five lab metrics: First Contentful Paint, Largest Contentful Paint, Total Blocking Time, Cumulative Layout Shift, and Speed Index. Each metric has a non-linear scoring curve. A 0.1-second improvement on LCP at 2.5s is worth more points than the same improvement at 5s.

Two consecutive PSI runs on the same store can differ by 10 to 15 points without you touching anything. The mechanism, documented in a Shopify Community thread from 2023, is that jquery.min.js served from Shopify's own CDN sometimes loads in 300ms and sometimes takes 800ms depending on whether your test request hit a warm or cold edge cache. That 500ms swing alone is enough to move PSI substantially. Apps that depend on jQuery inherit the variance. So does anything render-blocked behind it.

In practice: a stable PSI number does not exist for a Shopify store. Plus or minus 10 points between consecutive runs is the floor of normal variance. If you optimized something and your score went up by 8 points, you cannot tell from the score whether the optimization worked or whether the second run got lucky with a warm CDN edge.

This is why you use PSI for the Opportunities tab, not the headline number. Properly Size Images, Reduce Unused JavaScript, Eliminate Render-Blocking Resources - each item comes with a measured millisecond savings against this specific simulated run. Those are deterministic enough to act on. The score on top of the report is noise around a useful diagnostic.

What GTmetrix and Pingdom are doing differently

GTmetrix tests from a different city, with different default device settings, and reports a different blend of metrics. Out of the box it tests from Vancouver on a desktop profile with no throttling, which is roughly the opposite of a mid-range Android on cell. A "B" from GTmetrix and a 41 from PSI do not contradict each other. They are two different verdicts from two different judges on two different criteria.

Pingdom defaults to a tier-one US data center on an unthrottled network. It almost always returns a kinder number than PSI for the same store, because it is testing what an enterprise user in a US office would experience, not what a phone customer in Liverpool sees. Pingdom is fine for a desktop-skewed store with US-heavy traffic. It tells you very little about mobile performance, which is where Shopify checkout conversion actually lives.

Neither tool integrates with CrUX. Both report on a single synthetic load. That is why neither is the tool that ranks you in Google.

What CrUX is and why Search Console is your real benchmark

The Chrome User Experience Report is a public dataset of real-user performance metrics collected from Chrome users who have opted in to anonymous performance reporting. For each origin in the dataset, Google publishes p75 values across LCP, INP, CLS, and a few other metrics, aggregated over a rolling 28 days. "p75" means 75% of real user experiences were at least this fast. The 25% worst are what the metric is reporting.

Search Console's Core Web Vitals report surfaces your store's CrUX data, grouped by URL pattern. The Shopify Web Performance Dashboard surfaces the same numbers inside the admin without a Search Console hop, on stores with enough Chrome traffic for CrUX to build a profile. If your dashboard says "Not enough data," your store has not had enough recent Chrome traffic to populate the report. Brand-new stores often see this; established stores rarely do.

Two things to internalize about CrUX:

  1. It cannot be tricked by user-agent sniffing. The user agents in the report are real users on real Chromes. A cloaking script that detects "Chrome-Lighthouse" will not detect a real customer, so a real customer's data lands in your CrUX whether you want it to or not. This is why CrUX is the only number a cloaking gig cannot game.

  2. It updates on a 28-day rolling window. A fix you ship today shows up in the dashboard in three to four weeks. The frustration of waiting is the price of measuring real users instead of bots. If your client needs the number to move on the day of a push, they need a different number, and PSI is the only one that moves that fast - which is part of why merchants over-trust it.

How to use each tool, on purpose

Here is the assignment that has worked best for me across a few hundred audits.

QuestionTool
Will Google rank this page worse because it is slow?Search Console > Core Web Vitals
Did my fix actually move real users in the last 28 days?Search Console or Shopify Web Performance Dashboard
What is the largest opportunity I have not addressed yet?PSI Opportunities tab
Why does my page paint slowly on a slow phone, specifically?Chrome DevTools > Lighthouse, Performance tab traces
Is this single change I shipped half an hour ago measurably better?DevTools Lighthouse, 3+ runs, take the median
Does my store feel fast to me on my office desktop?Pingdom (and recognize that the answer is meaningless for mobile)

The headline scores on PSI, GTmetrix, and Pingdom go in the "do not put this in front of a client" pile. They move around too much, they reward gameable optimizations, and they are not the number that ranks you. Their underlying diagnostic data is useful. Their composite scores are not.

The store that needed this, end to end

The merchant who sent me the six screenshots had a hero image lazy-loading by default and Microsoft Clarity loaded synchronously in the head. Both showed up in the PSI Opportunities tab with concrete millisecond impact. We changed the hero <img> to eager loading and added a <link rel="preload"> for the same source. For the standalone version of the hero fix, including the preload trap, see the hero image LCP trap post, which walks through the lazy-loading default in detail. We deferred Clarity behind a small script that loads it after the first user interaction.

On the next PSI run, his score went from 41 to 67. He texted me an excited screenshot. I told him to ignore the screenshot.

The thing to watch was Search Console. Twenty-eight days later, his LCP p75 had moved from 4.7s to 2.9s. The "Poor" URL group disappeared. The CrUX dashboard moved from red to amber. That was the only piece of data I would have used as evidence in front of a client paying for the work. The PSI screenshot was a leading indicator, not a result.

If you have a real performance gap and you want a second pair of eyes on which tools to trust for which question, a free 15-minute chat with a help1 expert is enough time to walk through your specific store and the data each tool is showing for it. We are not going to send you a screenshot.

The one sentence

A lab test is a diagnostic. A field test is a verdict. Build your shopify page speed workflow around that distinction and the six different numbers stop bothering you. The only one that ranks you in Google is the time your real customers actually spent waiting for the hero image to render.

See also

Still stuck? Talk to an expert.

Our vetted Shopify experts can fix this issue for you in a live session. $39 per session. Your first 15 minutes are free.