
RETENTION MARKETING
Zero-Party Data: How DTC Brands Collect It and Turn It Into Revenue




Written & peer reviewed by Darkroom leardership
Last update: August 7, 2026
Zero-party data is information a customer intentionally gives you: preferences, intent, and context they choose to declare, usually through a quiz, survey or preference center. Unlike behavioral data you observe and interpret, it tells you why someone buys, which is the input personalization and AI systems can actually act on.
Two things you should know before reading anything else on this topic. Google kept third-party cookies, formally, in April 2025. And zero-party data became more useful anyway, for a better reason than the one most articles still give you.
What is zero-party data?
Zero-party data is anything a customer tells you about themselves on purpose: the skin type they select in a quiz, the categories they tick in a preference center, the “gift for my husband” they type into a survey.
The term is Forrester’s, defined by Khatibloo and colleagues in 2017 to separate what customers tell you from what you observe. Analysts also call it declared data, and the declaration is the point. Nobody inferred it. The customer said it.
That makes it different in kind, not just in degree, from the data most direct-to-consumer (DTC) brands run on. Purchase history and browsing behavior are records of what happened. A declaration is a statement of what the customer wants to happen, which is a far better instruction for retention marketing systems to follow.
The practical test is whether you had to guess. If a field in your customer profile was populated by a rule, a model or a pixel, it is observed. If it was populated because someone chose an option or typed an answer, it is declared, and you can act on it without hedging.
At Darkroom, a retention marketing agency serving high-growth consumer, mid-market and enterprise brands, this is where personalization programs start, because it is the only data type where the customer has told you the answer instead of leaving you to guess it.
Why did the case for zero-party data change in 2025?
Because the reason everyone gave for collecting it stopped being true, and a better reason took its place. For five years, this topic was sold as insurance against the death of third-party cookies. Then the death was called off.
On April 22, 2025, Google announced it would “maintain our current approach to offering users third-party cookie choice in Chrome, and will not be rolling out a new standalone prompt”. On October 17, 2025, it went further and retired most Privacy Sandbox technologies, including Topics and Protected Audience, citing low adoption.
Even the mobile side moved the opposite way from the forecasts: Apple App Tracking Transparency (ATT) opt-in had risen to roughly 50% globally as of 2025, up about 10 points since its 2021 launch, per AppsFlyer.
Some signal loss is still real, since Safari and Firefox block or partition third-party cookies by default and your measurement stack has to account for it. But the apocalypse this content category was built on did not arrive.
So the honest case is no longer defensive. It is this: observed data tells you what happened, declared data tells you why it happened, and only the second one is an instruction a personalization engine can follow with confidence. Everything below is built on that premise.

Zero-party data vs first-party data: what is the difference?
First-party data is what you observe about customers through their behavior; zero-party data is what customers deliberately declare to you. You own both, and both survived 2025 intact. The difference is reliability of meaning: behavior requires interpretation, declaration does not.
That distinction matters more than the four-way taxonomy marketers usually memorize, because it decides how much confidence a downstream system can place in a field. The full picture, from most reliable to least:
Data type | How you get it | Example | Reliability of meaning |
|---|---|---|---|
Zero-party | Customer declares it | “My skin is sensitive, I buy for two kids” | Highest: no inference needed |
First-party | You observe it on your properties | Purchase history, browse and cart behavior | High on facts, low on intent |
Second-party | A partner shares their first-party data | A retail partner’s audience overlap | Depends entirely on the partner |
Third-party | Bought or aggregated from external trackers | Interest audiences from data brokers | Lowest, and shrinking by browser policy |
The confusion between the first two rows costs real money. Take a beauty catalog: your systems can observe that a customer bought retinol twice. Only a declaration reveals it was a gift both times, for someone whose skin is nothing like theirs.
Every personalization decision built on the observed version of that customer is confidently wrong, and the error compounds with every send.
How do DTC brands collect declared data?
You collect it by asking, at moments where answering visibly improves the customer’s experience, and only ever asking for what you have already decided how to activate. That last clause is what separates a working collection program from the graveyard of unused custom properties.
Collection surface | What it captures | When to ask | Where it activates |
|---|---|---|---|
Product recommendation quiz | Needs, constraints, use case | Discovery, pre-purchase | Product recs, segment assignment, welcome path |
Email preference center | Categories, cadence, channel | Signup and ongoing | Send frequency, content blocks, suppression |
Post purchase survey | Buying-for, occasion, discovery source | Order confirmation window | Post-purchase branching, gifting flows |
Welcome-flow question | One high-value attribute | First email or SMS | Flow branching from message two onward |
Account and profile fields | Sizes, household, restock cycle | Account creation and reorder | Replenishment timing, size-in-stock alerts |
The stack matters less than the wiring. Whether you are running this inside Klaviyo or a dedicated customer data platform (CDP), every field in the table above should map to a named segment or flow branch before the form ships.
One rule prevents most of the waste: no new question ships without an owner and an activation. If nobody can name the segment it feeds and the flow that reads it, the field is a liability, because you have asked a customer for something and given nothing back.
What makes a product recommendation quiz worth building?
A quiz earns its build cost when every answer changes what the customer sees next, and it fails when it only changes what you store. The pattern that works, and the one we build: three to five questions, each answer mapped to a segment before launch, results that alter the product page, the welcome path and the replenishment logic.
The failure mode is just as specific. A quiz that collects skin type and then sends every respondent the same newsletter has not built a data asset. It has built a broken promise, and the customer noticed.
What is progressive profiling, and why does it beat one big form?
Progressive profiling means asking one question per interaction across the customer lifecycle instead of fourteen fields at signup. Completion rates decide the argument: nobody abandons a single well-timed question, and almost everybody abandons a long form.
The sequence writes itself from your activation map. Channel preference at signup, use case in the welcome flow, occasion in the post-purchase survey, household details at the second order. Within a quarter, you hold a profile no tracking pixel could have assembled, given voluntarily, one answer at a time.
How do you turn declared preferences into revenue?
Declared preferences earn revenue at three layers: segmentation, flow branching, and content blocks. Data sitting in a custom property earns nothing; every field you collect should be working in at least one of the three.
Segmentation turns declarations into audiences: the sensitive-skin segment, the gift-buyer segment, the buys-for-the-household segment. Flow branching splits journeys on what customers told you, so the welcome and post-purchase paths diverge by declared use case. Content blocks key dynamic modules to preference, which is what separates ecommerce personalization from first-name tokens.
Declared data also pairs with behavioral scoring rather than replacing it. RFM analysis, which scores recency, frequency and monetary value, tells you who is worth talking to; declaration tells you what to say. Run together inside a coherent email revenue architecture, the two compound customer lifetime value in a way neither achieves alone.
The effect is measurable. The personalization program behind Drip Hydration, a national health franchise, lifted conversion rates 42% year over year.
Measure it the way a retention marketing agency would: revenue per recipient and repeat rate against flow-level benchmarks, plus the customer retention metrics that read behavior change, never open rates.
The demand side is settled. McKinsey’s research, published November 2021 and still the standard citation, found 71% of consumers expect personalized interactions; a figure that old is a floor, not news.
How does declared data power AI personalization?
AI personalization is only as good as its inputs, and declared preference is the one input that corrects a model instead of flattering it. Train a recommendation system on browse behavior alone and it reproduces your merchandising bias, showing people more of what your site already pushed at them.
Feed it declarations and it learns what customers actually wanted, including everything they searched for and did not find.
The second mechanic is newer. As AI assistants take over product discovery, the traffic they refer arrives context-free: no keyword, no cookie trail, no session path to infer from.
Yet the traffic converts. Ahrefs found AI search visitors converting 23 times better than organic on its own site in June 2025, though that is one B2B software company measuring free-trial signups, not a retail benchmark.
The number that applies to consumer brands is Adobe’s, and it moved fast. Through the 2025 holiday season, AI referrals to US retail sites converted 31% better than other traffic, reversing a deficit that ran 23% the previous July. When an assistant sends someone already sold, asking one question is the only way to know who they are.
That is why the data layer and generative engine optimization are one system now: getting cited by AI assistants through AI search optimization brings the buyer, and declared data personalizes what happens after arrival. It is the architecture we built for Cocolab, whose site experience was rebuilt for both human buyers and AI models.
Where to start: your first 30 days
Start with one question, one segment, and one flow branch, then expand only after the loop closes. Four moves cover the first month:
Week 1: Pick the single highest-value attribute you cannot observe (use case, buying-for, restock cycle) and define the segment it will feed.
Week 2: Add one question capturing it to your welcome flow and post-purchase survey. One, not five.
Week 3: Branch one flow on the answer, with visibly different content per path.
Week 4: Read revenue per recipient by branch against your control, then decide the second question.
The 30-day constraint is deliberate. Programs that start with a fourteen-field ambition ship nothing; programs that close one loop earn the mandate for the next one.

Quick answers on declared data
Does asking hurt conversion? Asked at the right moment with a visible payoff, it lifts conversion; quizzes and preference prompts are experiences, not friction.
Do you need a CDP first? No. A well-configured email platform with custom properties activates declared data fine at most brand sizes.
Who should own it? Whoever owns lifecycle revenue. Data collected by a team that does not activate it goes stale in a quarter.
Work with a team that builds the data layer and activates it
If your personalization roadmap stalls at first-name tokens, the missing piece is rarely software. It is declared data wired into segments, flows and content, and a team that has done the wiring before.
Darkroom is the innovation agency high-growth consumer, mid-market and enterprise brands choose for high-profile go-to-market launches, and retention systems are where those launches keep earning.
Lifecycle programs live in 30 days, inside your existing Klaviyo, Attentive or Postscript stack
Program-level proof: an 85% lift in customer lifetime value and 50% one-year revenue growth behind Drip Hydration, published on our retention page
Collection surfaces, segments and flows built as one system, measured on repeat revenue
Talk to our retention team about building your zero-party data layer, and get the first loop closed inside a month.
Frequently Asked Questions
What is zero-party data?
Zero-party data is information customers deliberately share with a brand: preferences, intentions, context and constraints, given through quizzes, surveys and preference centers. It differs from data you observe or infer because the customer supplied the meaning directly, which makes it the most reliable input personalization systems can use.
What is the difference between zero-party and first-party data?
First-party data is behavior you observe on your own properties, like purchases and browsing. Zero-party data is what customers state outright, like their use case or preferences. Observation tells you what happened and leaves you to interpret it; declaration removes the interpretation step entirely.
Is zero-party data still worth collecting now that Google kept third-party cookies?
Yes, and for a stronger reason than the old one. The 2025 reversal killed the insurance argument, but declared data was never really about cookies: it captures intent that no tracking can observe, and it is the input AI personalization depends on. The case improved when the panic ended.
What are some examples?
Quiz answers about needs or constraints, preference-center selections for categories and cadence, post-purchase survey responses like occasion or buying-for, declared sizes and household details, and replenishment timing a customer sets. Each one is a statement the customer chose to make rather than a signal you inferred.
How do you collect it without hurting conversion?
Ask at moments where answering visibly improves the experience: a quiz that changes recommendations, a preference that changes cadence, a survey after purchase when attention is free. Keep it to one question per interaction, and never collect a field you have not already mapped to an activation.
Is declared data compliant with GDPR and CCPA?
Consent is explicit by construction, since the customer typed or selected the information knowingly. You still owe purpose limitation, a stated retention period, and a way to view, edit or delete what was shared. Treat those as design requirements and have counsel review the specifics; this is not legal advice.
How does it improve AI personalization?
It supplies the ground truth models cannot infer. Systems trained only on behavior amplify whatever your merchandising already showed people; declared preferences correct that bias with what customers actually wanted. And for AI-referred visitors who arrive with no trackable history, a single declared answer is the entire personalization input.





























































































































































































































































































































