Ranking & recommendation quality testing

Is your next launch actually making things better?

Know what improved. Know what got worse. Get direct insights on how to make it better. Know when to ship.

Whether you sell products, connect buyers and sellers, or serve content to a community, algorithms decide what people see. KeyBay Data tests any new algorithm, model or product change against your current experience, shows what got better and what regressed, and tells you whether it’s ready to ship.

One platform. Every surface.

  • Search resultsAre people finding what they came for?
  • RecommendationsAre “more like this” and “you might also like” getting better?
  • Browse and category pagesIs the best inventory rising to the top?
  • Home and discovery feedsIs the feed showing more of what each person cares about?
  • Emails and notificationsAre personalized messages surfacing the right items, posts or offers?
  • Sponsored placementsAre promoted results relevant without hurting the experience?
See how it works

Built for retailers, marketplaces, social and content platforms, and any team shipping ranking or personalization.

From “is it better?” to a clear answer in five steps

Here is what we do, using a made-up example. Keep scrolling and watch it play out.

1Choose the searches

Start with real searches from real people

We pick a sample of searches from your logs and sort them by type, so short lookups, long questions and messy typing all get tested.

2Run both versions

Ask the old and new search the same questions

Every search runs through your current algorithm and the new one at the same time. We save the top results from each, side by side.

3Rate every result

Custom rubrics and LLM-assisted grading, like a careful reviewer

We write scoring rubrics for your kind of search, covering several attributes such as relevance, completeness and freshness. An LLM then helps grade every result against them, the way a careful reviewer would, without knowing which version produced it.

4Read the scorecard

See where it improved and where it got worse

One overall number can hide problems. So the scorecard breaks results out by type of search and flags every group that moved up or down.

5Fix it, then decide

Get a plain go or no-go, with a fix list

For every group that got worse, you get a likely cause and a tip to fix it. Then we give you one clear decision on whether the new algorithm is ready to ship.

Step 1 of 5Choose the searches

A sample of 1,000 searches

Short and exact200invoice 48213
Long and wordy200how do I cancel without losing my data
Vague200something good to read on a long flight
Rare and niche200left handed violin teacher
Misspelled200accomodation policy

One search, two sets of results

how do I cancel without losing my data
Current A
  1. Pricing and plans
  2. Cancel your subscription
  3. Billing questions
  4. Export your data
  5. Account security
New B
  1. Cancel and keep your data
  2. Export your data
  3. Cancel your subscription
  4. Pause instead of cancel
  5. Delete your account

Same search, now graded

Graded onRelevanceCompletenessFreshness
Current A
  1. Pricing and plans○ Poor
  2. Cancel your subscription◐ Okay
  3. Billing questions○ Poor
  4. Export your data● Great
  5. Account security○ Poor

1 Great, 1 Okay, 3 Poor

New B
  1. Cancel and keep your data● Great
  2. Export your data● Great
  3. Cancel your subscription◐ Okay
  4. Pause instead of cancel◐ Okay
  5. Delete your account○ Poor

2 Great, 2 Okay, 1 Poor

Repeat for all 1,000 searches, and every group gets a score out of 100.

Scorecard: current vs new

71 → 74
Overall, a small gain. But one group got clearly worse, and that is hidden in the average.
Vague▲ Better +14
Current
57New
71
Long and wordy▲ Better +13
Current
64New
77
Short and exact● Same +1
Current
90New
91
Rare and niche● Same +2
Current
66New
68
Misspelled▼ Worse −13
Current
78New
65

What to fix

Misspelled searches▼ −13

TipThe new version stopped correcting typos. Turn spell correction back on before it ranks results.

The decision

Not ready to ship yet

Big wins on vague and long searches, but misspelled searches got worse. Apply the fix, re-run the same 1,000 searches, and it turns into a GO.

Our rule: ship only when no group of searches gets noticeably worse.

Ship search changes with confidence

Bring us your current search and the one you want to launch. We will send back a scorecard like this one, built on your own searches.

Get your scorecard

Contact us

Tell us about your search

Share a few details and what you want to compare. We will reply from andrew@keybaydata.com with next steps for your scorecard.