Paid Advertising

How to Build Lookalike Audiences That Actually Perform

Lookalike audiences fail more often from a weak seed than from a weak platform algorithm — most of the improvement available is in what you feed the model, not how you configure it.


A lookalike built from your full customer list — every purchaser regardless of value, tenure, or fit — will find people who resemble your average customer, which sounds reasonable until you remember that “average” includes the one-time discount hunters and the trial abandoners sitting right alongside your best accounts. The algorithm doesn’t know which of your customers you’d clone a thousand more of; it just finds people statistically similar to whatever list you handed it, which means the seed list is doing almost all the strategic work, and most advertisers spend their optimization effort on the wrong side of that equation.

The Seed Determines the Ceiling, Not the Percentage Slider

Every major ad platform lets you choose how tightly or loosely the lookalike matches your seed — a 1% lookalike is the tightest match, expanding to 5% or 10% for broader reach with looser similarity. Advertisers spend a lot of energy tuning this slider while feeding it a mediocre seed, which is optimizing the wrong variable. A 1% lookalike built from a weak seed (your full customer list with no value or quality filtering) will still find you more of the same mediocre customer profile, just very precisely. A 3% lookalike built from a genuinely excellent seed (your top 20% of customers by LTV) will often outperform a tighter lookalike built from an unfiltered list, because the seed quality sets a ceiling the percentage setting can’t overcome.

This means the actual optimization work should happen before you ever touch the lookalike percentage: deciding exactly which customers belong in the seed. Get that right and the percentage setting becomes a genuinely minor tuning decision. Get it wrong and no amount of percentage tuning fixes a lookalike built from the wrong population.

Seed From Value, Not From Conversion Alone

The default seed most advertisers reach for is “everyone who purchased” or “everyone who signed up,” treating every conversion as equally valuable input. This throws away the single most useful signal available: which of those customers were actually good customers. A seed built instead from your top-value cohort — top 20-25% by LTV, or by a proxy like retained past 90 days, or expanded to a higher plan tier — teaches the algorithm to find people who resemble your best outcomes, not your median outcome.

Building this seed requires connecting ad platform audiences to actual downstream value data, which is more setup than a simple “all purchasers” export but is close to the highest-leverage change available in lookalike strategy. If your CRM or billing system can export a list of customers above a specific LTV or retention threshold, uploading that specific list as a custom audience and building lookalikes from it — rather than from the platform’s default purchase-event audience — routinely outperforms value-blind lookalikes by a wide margin, because it’s solving for the thing that actually matters (find me more high-value customers) instead of a weaker proxy (find me more people who converted once).

Recency and Size Both Matter, and They Trade Off Against Each Other

A seed audience needs enough size for the algorithm to find real statistical patterns — most platforms recommend at least 100-1,000 people minimum, with lookalike quality generally improving up to a few thousand, after which additional size adds diminishing returns to lookalike quality specifically (as opposed to reach, which keeps growing). But a seed that includes customers from three years ago, acquired when your product, pricing, and target market looked different, dilutes the pattern with an outdated customer profile that may no longer represent who you’re actually trying to reach today.

The practical balance: seed from customers acquired within the last 6-12 months where possible, even if that produces a smaller list, rather than reaching further back purely to hit a size threshold. If the last 6 months doesn’t provide enough volume to seed a reliable lookalike, that’s itself useful information — it usually means the highest-value segment (top 20% by LTV) needs to expand to top 30-40% temporarily to get sufficient recent volume, rather than reaching back in time and including customers whose profile may no longer reflect your current ideal customer.

Exclude the Wrong People From the Seed as Deliberately as You Include the Right Ones

A seed audience contaminated with the wrong customers — refunded purchases, customers acquired through a since-discontinued discount promotion who never bought again, free-tier users who never converted — actively teaches the algorithm to find more people like that unwanted profile, right alongside the good signal from your genuine best customers. This dilution is invisible in the lookalike audience itself (you never see who specifically got matched), which is exactly why it goes uncorrected for so long — the lookalike just quietly underperforms and nobody traces it back to seed contamination.

Building an explicit exclusion pass into the seed-building process — pulling refunded customers, known low-value segments, and promotion-only buyers out of the seed list before uploading it — is a small amount of extra data work that measurably improves lookalike quality, particularly for businesses where a meaningful share of purchases come through heavy discounting or a specific low-quality acquisition channel that shouldn’t be setting the pattern for future targeting.

A Worked Example: What the Seed Math Actually Looks Like

Concrete numbers make the seed-quality argument easier to apply. Say a company has 2,400 total customers acquired over the past two years. The instinctive move is uploading all 2,400 as the seed. Instead, pull LTV data and rank them: the top 20% by LTV is 480 customers, and cross-referencing against acquisition date narrows that further to customers acquired in the last 12 months, landing at roughly 310 people — below the 1,000-person minimum most platforms suggest for a reliably stable lookalike.

This is the exact tension the earlier trade-off describes, and the resolution is the same: rather than reaching back to years-old customers to pad the list to 1,000+, expand the value threshold from top 20% to top 35%, which brings the 12-month-recent pool up to roughly 840 customers — still under the platform’s stated minimum, but usable, and platforms generally still build a workable lookalike below their stated ideal, just with somewhat lower match confidence. A/B testing this 840-person, top-35%, 12-month-recent seed against a 2,400-person, unfiltered, 24-month seed at equal budget over four weeks is the actual test of whether the quality trade-off was worth it. In cases like this, the value-filtered seed with fewer people routinely produces lower CPA and, more importantly, higher 90-day retention among the customers acquired through it, confirming that seed quality outweighed the size shortfall.

The number that matters most from this exercise isn’t the specific thresholds — those vary by business — it’s the discipline of running the actual count before deciding on a strategy, rather than assuming either “more data is always better” or “our top segment is too small to use,” both of which are guesses until you’ve actually pulled the numbers.

When You Don’t Have Enough Customers to Seed a Lookalike at All

Early-stage companies and businesses with genuinely small customer bases run into a real version of the size problem above: even the full, unfiltered customer list doesn’t clear a platform’s recommended minimum, let alone a value-filtered subset of it. Waiting until you have “enough” customers to build a proper lookalike often means waiting months while running blind on broad interest targeting in the meantime, which is its own cost.

The practical workaround is seeding from a proxy audience that’s larger than the direct customer list but still correlated with the customers you actually want. Options in rough order of usefulness: an engaged-email-subscriber list filtered to people who’ve opened or clicked in the last 90 days, which is usually larger than the paying-customer list and still signals real interest; high-intent website visitors, such as people who viewed pricing or a demo page multiple times, pulled from your analytics platform’s audience export rather than the ad platform’s own weaker “all site visitors” pixel audience; or, if you have any early customers at all, layering a small direct-customer seed with one of these larger proxy audiences using the platform’s audience combination tools, so the lookalike leans on the proxy for volume while still being anchored to actual paying customers rather than floating free of any real conversion signal. None of these fully substitutes for a mature, value-filtered customer seed, but each gets you a usable lookalike months earlier than waiting for the customer list itself to grow into a workable size.

Build Separate Lookalikes for Separate Customer Segments Rather Than One Universal Seed

A business selling to multiple distinct customer segments — different company sizes, different use cases, different price points — that builds one universal lookalike from its entire customer base is asking the algorithm to find “the average of several different things,” which produces a blurrier, less useful pattern than segment-specific lookalikes would. A B2B tool selling to both solo freelancers and 50-person agencies has fundamentally different ideal customer profiles for each segment, and blending them into one seed teaches the algorithm neither pattern clearly.

Splitting the seed by segment — a lookalike built specifically from your best agency customers, run in its own campaign, alongside a separate lookalike built from your best freelancer customers, run in a separate campaign with its own messaging — lets each lookalike find a sharper, more distinct pattern, and lets you write ad creative specific to each segment rather than generic creative trying to speak to both audiences at once. This does mean more campaigns to manage, but the quality improvement in targeting alone usually justifies the added management overhead for any business with genuinely distinct segments.

Refresh the Seed on a Cadence, Not Once and Forget It

A lookalike seed built once at campaign launch and never refreshed slowly drifts out of alignment with your actual current best customer profile as the business, product, and market evolve. A company that’s shifted upmarket over the past year but is still seeding lookalikes from a customer list heavy with its earlier, smaller-account customer base is targeting toward yesterday’s ideal customer, not today’s, and the mismatch grows larger the longer the seed goes unrefreshed.

Setting a refresh cadence — quarterly is reasonable for most businesses, monthly for fast-growing or fast-pivoting ones — where the seed list is rebuilt from the most recent qualifying cohort keeps the lookalike aligned with the current, not historical, ideal customer profile. This is a simple recurring task, but it’s frequently skipped because nobody owns it explicitly; assigning it as an actual recurring line item in the paid media team’s calendar, rather than an ad hoc “we should probably redo that at some point” task, is often the difference between a lookalike strategy that stays sharp and one that quietly decays over a year of campaign inertia.

The Failure Mode Where Your Own Lookalikes Compete With Each Other

Running multiple lookalikes at once — say a 1% and a 3% from the same seed, or segment-specific lookalikes for agencies and freelancers as described above — creates a real risk of audience overlap, where the same person qualifies for more than one of your active audiences and your own campaigns end up bidding against each other in the platform’s auction. This is invisible in the reporting for any single campaign, since each one just shows its own performance, but it quietly inflates your effective CPMs across the account and skews which campaign gets credit for a conversion that either one could plausibly have driven.

Most platforms provide an audience overlap tool for exactly this reason — check it whenever you’re running more than one lookalike simultaneously, especially when they share a seed or a similarity percentage close enough to substantially overlap (a 1% and 3% lookalike from the same seed will always have meaningful overlap, since the 3% audience is a superset that contains the 1%). Where overlap is high, either add mutual exclusions between the campaigns so a given person only falls into one audience at a time, or consolidate into a single campaign with the broader audience rather than running both and letting them cannibalize each other’s delivery. Segment-specific lookalikes built from genuinely distinct populations (agencies versus freelancers, as in the earlier example) typically overlap far less than percentage variants of the same seed, but it’s still worth a periodic check rather than an assumption.

Validate Lookalike Performance Against a Broad Interest-Based Control

The only way to know whether the effort invested in seed quality is actually paying off is to run the refined lookalike against a genuine comparison — typically a broad, interest- or demographic-based targeting campaign with similar budget and creative, run concurrently over the same 3-4 week window. Comparing cost per acquisition and, more importantly, downstream value or retention of the customers acquired through each targeting approach (not just the initial conversion cost) tells you whether the lookalike is actually finding better customers, not just cheaper conversions that turn out to be lower-value once they’re through the door.

This downstream check matters because lookalikes optimized purely for conversion cost can still find customers who convert cheaply but churn fast or never reach meaningful LTV — the whole point of seeding from value rather than raw conversions was to avoid exactly that outcome, so validating on the same value metric the seed was built from (not just cost per lead) closes the loop and confirms the strategy is actually working the way it was designed to.

Book a demo