What We Learned Building Personalized Storefronts for Seven Retailers
Back to blog

What We Learned Building Personalized Storefronts for Seven Retailers

Priya Nair 10 min read

In the second half of 2025, we ran our early-access program with seven retailers across four product verticals: home goods, outdoor apparel, artisan food and pantry, and specialty beauty. The program ran for six months from initial integration through a full holiday period for the retailers who needed seasonal coverage. Here is what we expected going in, what the data actually showed, and the two things we had to completely rethink.

This is a practitioner retrospective, not a marketing piece. Some things worked better than expected. One category behaved completely differently from our model. And we made two integration decisions early on that we would change if we ran the program again today.

The Setup

All seven retailers were mid-size independent operations: online-only or online-primary businesses with monthly shopper sessions between 8,000 and 60,000, and repeat buyer bases ranging from 800 to 4,500 unique customers with at least two purchases in the prior 12 months. None were Shopify Plus or enterprise platform customers. They were the kind of retailers that are too large for the generic "add a widget" recommendation tools but too small to commission custom personalization engineering.

The integration work took between one and four days per retailer depending on how cleanly their purchase history data was structured. Two retailers required a data cleanup pass before we could connect their order history to our model in a usable way. This cleanup work was not part of our original timeline estimate and it became something we addressed later in the process. More on that below.

Each retailer ran a 60-day controlled period where the personalized category page sort was shown to a randomly selected 50% of sessions while the other 50% continued to see the original default sort. This gave us a clean baseline comparison for each retailer across their own traffic mix.

What the Data Had in Common

Across all seven retailers, the personalized sort produced a higher category add-to-cart rate for the repeat buyer segment compared to the control group seeing the default sort. The effect was consistent in direction if not in magnitude. The smallest improvement we saw was around 8% above control for one of the artisan food retailers whose repeat buyer base was already extremely loyal and whose default sort was actually quite good. The largest was above 30% for a home goods retailer whose default sort was alphabetical and whose repeat buyer preferences were concentrated in a few sub-categories that alphabetical sort consistently buried.

The finding that surprised us most was how quickly the improvement appeared. We expected the model to need several weeks to calibrate. In practice, by day 14, the direction of the effect was already visible. The model was working with purchase history that existed before we started, so it was not learning from scratch. It was applying pre-existing pattern recognition from day one, and the category pages were better from the first sessions.

The personalization effect was visible within the first two weeks for every retailer. The model was not learning from zero. It was reading patterns that had been sitting in purchase data for months, untouched.

Where the Verticals Diverged

The behavior we did not predict was in the specialty beauty vertical. We had two beauty retailers in the cohort. Both showed strong improvement in category add-to-cart rate overall. But the pattern underneath that improvement was different from the other verticals in a way that changed how we think about replenishment categories.

In home goods and outdoor apparel, shopper preferences showed gradual trajectory evolution: shoppers moved from entry-level items toward higher-quality or more specialized products over multiple purchase cycles. The personalization model works well in trajectory-oriented categories because purchase history creates a clear path to follow.

In beauty, shopper behavior was much more replenishment-oriented. Many shoppers were buying the same product repeatedly, sometimes the exact same SKU. The co-purchase patterns were less about trajectory and more about reliable re-ordering. In this context, the most valuable recommendation the model could make was often the same item or a direct replacement from the same line, not an adjacent product that fit a broader trajectory.

We had not built a strong replenishment-recency signal into the initial model. We added it after observing the beauty retailer data, and the improvement in recommendation accuracy for replenishment categories was meaningful. If we were starting over, replenishment detection would be part of the baseline model architecture, not an add-on.

Two Things We Had to Rebuild

The first was data quality handling. We assumed that mid-size retailers using standard e-commerce platforms would have consistently structured purchase history data that we could read directly. That was true for five of the seven. For two retailers, the data had inconsistencies: shopper identifiers that had changed between platform migrations, orders that had been imported from a legacy system with different SKU formats, return records that were stored separately from order records without a consistent join key.

This meant we were doing data detective work during the integration phase for those two retailers, which delayed their go-live dates and added stress to a timeline we had committed to. We now include a purchase history audit as a pre-integration step for every new merchant. Before we touch the category pages, we run a structured check on purchase data quality and confirm that the data we need is in a joinable state. This adds time upfront but eliminates the mid-integration surprises.

The second rebuild was in how we handled holiday and promotional periods. Two of our retailers ran significant promotional events during the program period. Their purchase data during those windows was noisy in a specific way: promotional buyers, who came in on a discount and bought items they would not have bought at full price, were generating purchase signals that did not accurately represent their stable preferences. We were learning from a temporary behavioral state.

We added a promotional flag layer that down-weights purchase signals from sessions that originated from promotional email campaigns or discount code entries. This keeps the model from treating a promotional purchase as evidence of a strong product preference when it may actually be evidence of a strong discount preference. The flag is configurable by the retailer so they can define what qualifies as a promotional session for their business.

What We Told the Retailers About Their Own Data

One of the more useful side effects of the early-access program was that every retailer came away knowing things about their shopper base they had not known before. The purchase pattern analysis we run during integration surfaces information that is not visible in standard platform analytics.

One outdoor apparel retailer discovered that their highest-lifetime-value shopper segment, which accounted for roughly 18% of buyers but over 35% of revenue, was concentrated almost entirely in two sub-categories that their category pages were consistently underweighting in the default sort. Their default sort was optimized for new arrivals and site-wide best sellers, both of which skewed toward a different segment. The high-value repeat buyers were effectively being ignored by the page presentation their store was designed to show.

Knowing that is useful even beyond the personalization output. It informs inventory decisions, email marketing segmentation, and where to invest in product development. The purchase pattern data contains retailer intelligence that is not specific to our product. We started including a shopper segment summary report as part of the integration deliverable because the retailers found it independently valuable.

The Metric That Mattered Most

The metric we cared most about going into the program was category add-to-cart rate for repeat buyers. That stayed central. But by the end of the program, the metric that the retailers themselves responded to most strongly was repeat buyer retention at 90 days: the share of repeat buyers who made another purchase within 90 days of their most recent order.

In hindsight this makes sense. Category add-to-cart rate is a session metric. It tells you whether a session that included a category page visit produced an add-to-cart action. Repeat buyer 90-day retention is a relationship metric. It tells you whether shoppers are staying engaged with the store over time.

The retailers saw the retention metric move more than they expected. The interpretation that made most sense to them: when repeat buyers are consistently shown relevant products on category pages, they have more positive sessions, which reinforces the habit of returning. The engagement was not just improving individual session conversion. It was strengthening the overall relationship.

Whether that retention improvement was caused by the personalization directly or by correlated changes during the program period is hard to isolate conclusively from a six-month observation window. We are running a longer observation on a subset of the original cohort to get cleaner signal on the retention question. What we can say is that the direction was positive across all seven retailers, and the retailers reported that their repeat buyer base felt more active during the program period than in comparable prior periods.

See how Curated For You rebuilds your storefront.

Book a demo Back to blog