
A Lighthouse score is a lab result from one simulated device. Google assesses Core Web Vitals on field data at the 75th percentile of real page loads, segmented by mobile and desktop. A Hyvä migration that moves your lab score from 32 to 94 and your field data barely at all has not delivered what you paid for. Baseline the field numbers before you start.
This distinction gets blurred constantly in migration case studies, usually not dishonestly. Lab scores are easy to capture, they move dramatically, and they make a good slide. Field data moves slower, is harder to attribute, and is the only thing that affects rankings and revenue. If you are about to commission or justify a frontend migration, this is the measurement discipline that protects you.
The three metrics and what “good” means
Per Google’s Core Web Vitals documentation:
| Metric | What it measures | Good threshold |
|---|---|---|
| Largest Contentful Paint (LCP) | Loading | Within 2.5 seconds |
| Interaction to Next Paint (INP) | Interactivity | 200 milliseconds or less |
| Cumulative Layout Shift (CLS) | Visual stability | 0.1 or less |
Two details in that documentation matter more than the numbers themselves.
The assessment point is the 75th percentile of page loads, segmented across mobile and desktop. Not the median, not the average. That means a quarter of your visits can be slow and you still pass, but equally, an average that looks healthy can hide a failing 75th percentile. Any reporting that quotes an average is answering a different question than Google is asking.
And Google is explicit that lab measurement is not a substitute for field measurement. Device capability, network conditions, and real user interaction patterns are things a simulated test cannot reproduce.
Why Hyvä migrations flatter lab scores
Hyvä removes a large amount of JavaScript. Lab tools measure a cold load on a simulated device, which is precisely the scenario where removing JavaScript produces the most dramatic-looking improvement. So the lab number moves hard and immediately.
Field data moves differently, for reasons worth understanding rather than being surprised by.
It lags. Field datasets are rolling collections over a trailing window. Your migration does not appear as a step change on the day you deploy. It appears gradually, and a report pulled a week after launch is still mostly measuring the old site.
It includes everything you did not change. A migration replaces the frontend. It does not fix a slow time to first byte, an overloaded database, an unoptimised image pipeline, or a third-party tag manager loading eleven scripts. If your LCP was dominated by server response time, Hyvä will disappoint you, and that is not Hyvä’s fault.
It reflects your actual traffic. Your users’ devices and networks, not a simulated mid-tier phone. A store with a largely mobile audience on variable connections sees a different result from one serving desktop office traffic.
Interactivity is where the real win usually is. INP is about responsiveness to interaction, which is exactly what a heavy JavaScript stack degrades. In our experience the honest headline for a Luma to Hyvä migration is more often an INP story than an LCP one, even though LCP is what gets quoted.
What to baseline before you touch anything
This is the part teams skip and then regret, because once the old frontend is gone you cannot go back and measure it.
Capture field data per template, not sitewide. Homepage, category, product detail, cart, and checkout behave differently and a sitewide figure hides which one is failing. Use the Chrome User Experience Report and Search Console’s Core Web Vitals report, and record the 75th percentile for LCP, INP, and CLS on mobile and desktop separately.
Record the date and the window. Because field data is a trailing aggregate, a before-and-after comparison is meaningless without knowing what period each number covers.
Capture lab results too, but label them. PageSpeed Insights gives you both, and Lighthouse gives you the lab detail and the diagnostic opportunities. Lab data is genuinely useful for finding what to fix. It is just not the proof that you fixed it.
Write down what else is changing. If you are also moving hosting, upgrading Magento, or replacing a tag manager in the same window, say so now. Otherwise you will spend the post-launch review arguing about attribution with no way to settle it.
Note your server response time separately. If time to first byte is a large share of your LCP, decide before the project whether you are addressing it. This single check prevents the most common disappointed outcome.
Reading the result honestly
Wait a full data window after launch before drawing conclusions, and then ask three questions.
Did the 75th percentile move, on mobile, per template? That is the number. If mobile product detail LCP went from failing to passing, you have a result worth reporting.
Did INP improve more than LCP? Usually yes for a Hyvä migration, and that is a good sign the JavaScript reduction did what it was supposed to.
What did not move, and why? If LCP is stubborn, the cause is almost always server response time, image delivery, or a third-party script, none of which a frontend framework change addresses. This is useful information, not a failure. It tells you where the next piece of work is, and we set out that work in our guide to fixing Magento Core Web Vitals.
Be disciplined about attribution. If you migrated the frontend and changed hosting in the same month, you cannot cleanly credit either. That is an argument for sequencing changes when you care about measurement, and for saying plainly in your reporting that the numbers are combined when you did not.
The reporting template we use
A migration report that survives scrutiny fits on one page and has the same shape every time.
Start with a table of the five key templates down the side and six columns across: mobile LCP, INP, and CLS before, then the same three after. All at the 75th percentile, with the collection window for each stated in a footnote. Nothing else in the table. No lab scores, no averages, no percentages.
Underneath, three short paragraphs. What passed that previously failed, stated per template. What did not move, with the reason if you know it. And what else changed in the same window that could have contributed, listed plainly even when it weakens the story.
The desktop equivalent goes in an appendix. For most ecommerce merchants desktop passes comfortably before and after, so leading with it makes a migration look less effective than it was while burying the segment where the work actually landed.
The reason to fix the format in advance is that it removes the temptation to select the flattering metric after the fact. If everyone agrees before the project what the report will contain, nobody spends the review arguing about which number to lead with.
Turning measurement into a business case
Performance work gets funded when it is expressed in revenue rather than milliseconds, and the honest version of that model starts with your own baseline rather than an industry statistic.
The chain is straightforward: a measured improvement at the 75th percentile on your highest-traffic template, applied to your own conversion rate and average order value, over your own session volume. Every one of those inputs should come from your analytics, not from a case study about another merchant. We built out that model in detail in how to justify a Luma to Hyvä migration to your CFO.
The reason to insist on your own numbers is not purity. It is that a business case built on someone else’s conversion lift falls apart the first time a finance team asks where the figure came from, and it takes the credibility of the whole project with it.
A note on what performance work cannot buy
Worth saying plainly, because we would rather scope honestly than oversell. Passing Core Web Vitals is a threshold, not a ranking lever you can keep pulling. Going from failing to passing is valuable. Going from passing comfortably to passing very comfortably generally is not, and the effort is usually better spent elsewhere.
The same applies to lab scores. Chasing a Lighthouse score from 92 to 98 is optimising a number that no user experiences. If your field data passes on every template, the performance project is done and the next constraint on revenue is somewhere else.
How Bemeir helps
Bemeir is a Brooklyn ecommerce agency and the USA’s first official Hyvä Gold Partner. We baseline field data before a migration starts, because we would rather be measured on the number that matters than on a lab score that flatters everyone.
Our Hyvä development services cover the migration and the measurement around it, and our Magento and Adobe Commerce development practice covers the server, caching, and image work that a frontend change alone does not fix. We also build on Shopify and Shopify Plus, Shopware, and BigCommerce, and the CDN, image, and monitoring tooling behind a performance programme comes from our technology partner ecosystem.
More about the team is on the about Bemeir page, or start a conversation from the Bemeir homepage.
Frequently asked questions
What is the difference between lab and field data for Core Web Vitals?
Lab data comes from a controlled, simulated test on a defined device and network, such as Lighthouse. Field data comes from real users on their own devices and connections, collected in datasets like the Chrome User Experience Report. Google assesses Core Web Vitals on field data, and its own documentation states that lab measurement is not a substitute for field measurement. Lab data is for diagnosis; field data is for assessment.
What are the current Core Web Vitals thresholds?
Largest Contentful Paint should occur within 2.5 seconds, Interaction to Next Paint should be 200 milliseconds or less, and Cumulative Layout Shift should be 0.1 or less. Assessment uses the 75th percentile of page loads, segmented across mobile and desktop, so a quarter of visits can exceed the threshold and the page still passes.
Why did my Lighthouse score jump after a Hyvä migration but my field data barely move?
Most likely because the remaining bottleneck is something the frontend does not control. Slow server response time, unoptimised images, or heavy third-party scripts all cap your field performance regardless of how light the theme is. Field data also lags, since it is a trailing aggregate, so a report pulled shortly after launch is still largely measuring the old site.
How long should I wait before measuring a migration’s impact?
At least one full field data window after launch, and ideally longer, because the dataset is a rolling aggregate rather than a daily snapshot. Pulling a comparison a week after deployment will mix old and new traffic and understate the result. Record the date and window of both your before and after measurements so the comparison is meaningful.
Which Core Web Vital does a Hyvä migration improve most?
In practice, Interaction to Next Paint tends to show the clearest improvement, because INP measures responsiveness to user interaction and that is exactly what a heavy JavaScript stack degrades. LCP usually improves too, but it is more constrained by server response time and image delivery, which a frontend change does not address. Reporting that leads with LCP is often quoting the less honest number.





