We'd been working on the model for about six months when we connected with the retailer who would become our first pilot partner. A home goods store with a loyal customer base, running on Shopify, roughly 18,000 monthly visitors. Healthy repeat-purchase rate. A team that cared about their customers and felt, correctly, that their storefront wasn't doing those customers justice.
Their category pages showed the same top-20 products to everyone who visited. Alphabetical after the first page. No personalization logic at all. The store had been built competently but without any mechanism to adapt to what they'd learned about their customers over three years of sales history.
We went in confident we understood the problem. We came out knowing we'd understood it only partially.
What the Data Showed That We Didn't Expect
Before we touched anything, we spent three weeks just analyzing their purchase history. Three years of order data, roughly 40,000 transactions, across a product catalog of about 340 SKUs.
The first thing we found confirmed our hypothesis: repeat buyers had a very different purchase pattern from first-time buyers. Repeat buyers were buying from a narrower range of product attributes within categories. They had taste profiles. A customer who had bought linen bedding twice was not in the market for polyester bedding. A customer who consistently bought in a mid-to-premium price range was leaving the page quickly when the default sort put budget items front and center.
What we didn't expect: the range of taste profiles among repeat buyers was much wider than we'd assumed. We'd imagined a handful of coherent clusters. What we saw was something closer to a spectrum with dozens of distinguishable positions along it. Two customers who had both bought the same best-selling throw blanket had almost nothing else in common in their subsequent purchase behavior.
This mattered architecturally. Our initial model design was optimized for a small number of segments. Assign a shopper to a segment, serve that segment's preferred sort order. Clean, fast, interpretable. That design would have been adequate for five segments. It was not designed for 40 distinct taste profiles and the graceful interpolations between them.
We rebuilt the core scoring logic before we ran a single live session.
The Integration Reality
The integration process with the retailer's Shopify store took longer than we'd planned, for reasons that became a core design input for every integration since.
The purchase history data was clean. Three years of structured order data in Shopify is well-formatted and accessible. The product catalog metadata was less clean. SKU descriptions had evolved over the years. Category assignments were inconsistent in some cases: the same product type lived in two different categories depending on when it had been added to the catalog. Some product attributes that we needed for the scoring model (material type, size range, color family) were present for about 70% of the catalog and missing for the rest.
We spent time on data normalization that we hadn't budgeted for. This was useful, because it forced us to build normalization tooling that later became a standard part of onboarding. But it delayed the live pilot by two weeks.
The lesson: the quality of the personalization is bounded by the quality of the product metadata. Retailers who have invested in clean, consistent attribute tagging get better results faster. Retailers with inconsistent metadata get good results eventually, but the onboarding takes longer. We now surface this as a diagnostic during the initial integration conversation.
First Four Weeks of Live Data
We ran the pilot on three category pages: Bedding, Throws and Blankets, and Kitchen Textiles. These were the retailer's highest-traffic categories and the ones with the strongest repeat-buyer concentration.
We didn't do a full A/B split because the retailer wasn't set up for that technically and we didn't want to add another integration layer. Instead, we compared the four-week post-launch period to the same four weeks in the prior year, adjusted for a roughly 8% year-over-year traffic increase.
The results were encouraging in aggregate but more interesting at the segment level. Overall category add-to-cart rate improved by about 14 percentage points across the three categories combined. But the distribution was uneven in an instructive way.
Repeat buyers with five or more previous purchases saw the largest improvement. Their category add-to-cart rate increased by 26 percentage points. These were the shoppers with the richest purchase histories and the clearest taste profiles. The model had the most signal to work with.
Repeat buyers with one or two previous purchases saw a smaller but still meaningful improvement. First-time visitors saw almost no change in category conversion, which is what we expected: no purchase history, no personalization signal, same sort order as before.
This distribution confirmed something we'd suspected but hadn't proven: the value of purchase-history personalization scales with purchase history depth. The more a shopper has bought, the better the model can serve them. This isn't a limitation so much as a property that aligns naturally with retailer economics. Your most loyal shoppers are also your most valuable. Personalization works best exactly where the business value is highest.
What We Had to Rebuild After Week Two
Two weeks in, we noticed an anomaly. A small cohort of shoppers who had very long purchase histories (ten or more orders over three years) were seeing erratic sort orders. Products that didn't seem to fit their established taste profile were surfacing in the top positions for them.
The problem was in how we were handling purchase recency. Our initial scoring model weighted older purchases and newer purchases equally, treating all of a shopper's history as a flat signal. For shoppers with long histories, this was a mistake. A customer's buying behavior from three years ago isn't necessarily predictive of what they want today. Tastes change. Life circumstances change. A customer who bought children's bedding regularly three years ago may no longer have children in that size range.
We added a recency decay function to the scoring model. More recent purchases get higher weight. Purchases older than 18 months decay toward baseline. The anomalous sort orders resolved within a few days of deploying the fix.
This is a recurring theme in building anything on behavioral data: the edge cases that reveal architectural assumptions you didn't know you'd made. Long-tenure customers with evolving taste profiles exposed a recency problem we'd overlooked. It's the kind of thing you can only find with live data.
Six Months Later
The retailer is still using Curated For You. They've since added personalization to two additional category pages. The gains have been consistent, though not as dramatic as the initial burst: add-to-cart rate improvements tend to be largest in the first few weeks as the model converges on established shopper profiles, then stabilize as the personalization becomes the new baseline.
The most meaningful feedback from the retailer came about four months in, when they described their repeat buyers as "shopping differently." Customers were visiting fewer pages per session but adding to cart earlier. They were buying from parts of the catalog they hadn't visited before, because those items were surfacing earlier in the sort for them based on purchase pattern matching. The long tail of the catalog was getting more exposure to the customers most likely to buy from it.
That last point was something we hadn't explicitly designed for. It emerged from the model doing what it was built to do: surface what each shopper is most likely to buy. For a retailer with a deep catalog, that incidentally improves long-tail product discovery for the customers whose purchase history puts them in range of those products.
The first pilot taught us more than we expected. It confirmed the core thesis, surfaced three architectural decisions we had to revise, and set the standard for how we approach every new merchant integration now.