Purchase intent data is any signal indicating that someone — or an audience — is actively considering a purchase in a category right now: comparing models, checking prices, reading reviews. It has classically been harvested by tracking individuals. This guide explains the traditional sources, and a newer alternative: inferring the in-market state of a page's readership from the page's content itself — 283 intent segments in 34 groups, no cookies, no IDs, no tracking.
Purchase intent data captures the difference between someone who likes a category and someone who is about to buy in it. A person can be interested in cars for thirty years; they are in-market for a car for perhaps ninety days of that. Intent signals are the observable traces of that in-market window: the searches, comparisons, calculator sessions, spec-sheet downloads and review-reading binges that cluster tightly around a purchase decision and then stop.
That time-bounded quality is what makes intent data the most valuable targeting input in advertising. Interest data describes a durable affinity and is useful for reach and brand planning; intent data describes a transient state and is useful for conversion. An ad for mortgage rates shown to someone actively comparing lenders performs on a different order from the same ad shown to a general finance audience — which is why in-market segments have historically commanded the highest CPMs of any third-party data category, and why B2B intent platforms became a standard line item in enterprise marketing budgets.
The catch has always been sourcing. Most intent data was built by observing individuals — their queries, their browsing, their transactions — which means it inherited every problem of user-level tracking: cookie dependence, consent requirements, coverage that collapses on Safari, Firefox and iOS (roughly 40%+ of traffic carries no usable third-party cookie today, and third-party cookies persisting in Chrome does not change that), and growing regulatory exposure under GDPR, CCPA and their successors. This page is part of our hub on cookieless audience segmentation, and it looks at intent through that lens: what the classic sources actually measure, what a content-inferred alternative measures instead, and how the two compare on privacy, coverage, latency and granularity.
One definitional point worth fixing early: intent exists at different resolutions. It can be claimed about an individual ("this user is in-market for an SUV"), an account ("this company is researching data warehouses"), or an audience ("the readership of this page skews heavily toward people shopping for a new vehicle"). The three are often conflated in vendor marketing. They should not be — they have different evidentiary bases, different privacy footprints and different appropriate uses, as the rest of this guide makes explicit.
Five source families account for nearly all commercial intent data. Each observes a different behavior, at a different resolution, with different blind spots — understanding them is the fastest way to evaluate any intent product's claims.
Queries are the purest intent signal that exists: "best 7-seater EV 2026 price" is a person announcing their in-market state in their own words. Search engines monetize this directly through search ads, and query-derived audiences power much of the in-market segmentation inside the large ad platforms' walled gardens.
Limit: the raw signal stays inside the platforms that captured it; outside them, buyers get pre-packaged segments with little transparency into recency or construction.Product views, cart adds, configurator sessions, pricing-page visits — behavioral events collected on a marketer's own properties (first-party) or historically across the web via third-party cookies and SDKs. This is the signal behind retargeting and most "in-market" segments sold on data marketplaces.
Limit: cross-site collection depends on tracking infrastructure that is blocked by default on Safari, Firefox and iOS, and consent-gated everywhere else.Opted-in panels of users who share their browsing, search or purchase activity in exchange for compensation. Panels observe deep, longitudinal behavior for a small consented population, which is then statistically projected onto the broader market. Widely used for measurement, category sizing and calibrating other datasets.
Limit: panel sizes are small relative to the web, projection introduces modeling error, and niche categories may have too few in-panel buyers to read.Card networks, retailers, receipt-scanning apps and e-commerce platforms see actual purchases — the ground truth every other source approximates. Transaction-derived segments ("bought baby products in the last 90 days") are strong predictors of adjacent and repeat purchases.
Limit: it is inherently backward-looking — it tells you what was bought, not what is being considered — and it is among the most sensitive personal data categories under privacy law.B2B intent vendors watch for content-consumption spikes at the account level: when an unusual number of readers resolved to one company start consuming content about, say, container orchestration — across publisher co-op networks, bidstream data or the vendor's own media — that account is flagged as "surging" on the topic and pushed to sales and ABM teams.
Limit: IP-to-company resolution is noisy (VPNs, remote work), topic taxonomies are vendor-proprietary, and the observation network covers only a slice of relevant reading.The newest family, and the subject of the rest of this page: instead of observing people, analyze the page. What a URL is about reveals the likely in-market state of its readership — no individual is observed at all. It is the intent analogue of contextual targeting, upgraded from "what is this page about" to "what is this page's audience shopping for".
Limit: it describes audiences, not individuals — a constraint we treat honestly in the limitations section below.Consider a mortgage-calculator page. Nobody reads it for fun. Its audience is, almost by definition, people in-market for a mortgage — and that is knowable from the page alone, without observing a single visitor. The same logic holds across the web: a "best CRM for small business 2026" comparison is read by CRM buyers; a hotel-review roundup for Lisbon is read by people planning travel; an EV charging-cost explainer is read by people considering an electric vehicle. The content selects the audience.
Content-inferred intent systematizes this. A model reads the page — its topic, angle, funnel position, price signals, comparison structure — and outputs the purchase-intent segments its readership plausibly occupies, each mapped to a controlled vocabulary and carried with a banded confidence (low / medium / high). The unit of analysis is the page or domain, never the person. That single design choice is what makes the approach privacy-safe by construction: there is no personal data anywhere in the pipeline, so there is nothing to consent, store, hash or delete.
In our implementation, intent is one of five signal families (alongside demographics, life stage, B2B firmographics and personas) returned by the audience segmentation API for any URL in real time, and pre-computed across a 102M-domain dataset for planning and curation work at scale. Intent is only asserted where content actually supports it — a mortgage calculator earns a high-confidence real-estate finance signal; a general news homepage earns none rather than a guess.
Intent codes are stable identifiers (PI.travel.hotels_and_resorts), so segments survive model updates, joins and multi-vendor pipelines. The full branch is browsable on the audience segmentation taxonomy page.
Neither approach dominates the other on every axis — they answer different questions. The honest comparison looks like this.
| Dimension | Panel / behavioral intent (user-observed) | Content-inferred intent (page-observed) |
|---|---|---|
| Data basis | Observed actions of individuals or accounts: queries, browsing events, transactions, panel activity, IP-resolved content consumption. | The page itself: topic, funnel position, comparison structure, price signals. The audience is inferred from what the content selects for. |
| Privacy | Processes personal data; requires consent management, contracts and deletion workflows; exposure under GDPR/CCPA and sensitive-category rules. | No personal data processed at any stage; privacy-safe by construction. No consent dependency, no data-subject requests, no browser-policy risk. |
| Coverage | Bounded by the tracking footprint: collapses where third-party cookies are blocked (Safari, Firefox, iOS), thins with consent opt-outs, and panels cover a small projected sample. | Works on 100% of pages and traffic, including cookieless environments, new visitors and unconsented sessions — the signal lives in the content, not the browser. |
| Latency | Segments are built from accumulated observations; individual membership often lags behavior by hours to weeks, and stale membership decays silently. | Evaluated at request time per URL (or from the pre-computed domain dataset); a new page can carry intent signals the moment it is published. |
| Granularity | Individual or account level — genuinely per-person when the data is good, which is its core strength and its core privacy cost. | Audience level: the aggregate in-market skew of a page's readership. Precise about pages, deliberately silent about persons. |
| Best use | Retargeting, closed-loop measurement, sales triggers on named accounts, suppression of existing customers. | Prospecting reach, in-market contextual targeting, inventory curation, seller-defined audiences, account scoring by what a company's buyers read. |
In practice sophisticated buyers run both: user-observed intent where consented first-party relationships exist, and content-inferred intent to extend in-market reach across the growing share of traffic where user-level signals are unavailable. The deeper comparison of the two philosophies — observation versus inference — is covered in our guide to contextual vs behavioral targeting.
Intent data is only as useful as the category system behind it. Free-text model output ("seems interested in cars?") cannot be traded, joined or audited — controlled codes can.
Every intent value we return comes from the Purchase Intent branch of our controlled vocabulary: 34 tier-1 groups containing 283 tier-2 in-market segments, each with a stable machine-readable code in the PI.* namespace and a human-readable label. The branch covers the complete IAB Purchase Intent structure — the vocabulary is versioned (v1.0) and aligned with the IAB Tech Lab Audience Taxonomy 1.1, which means segments built on it slot directly into industry pipelines that speak IAB, from data-marketplace listings to Seller Defined Audiences. How that alignment works across all our attribute families — and why standard codes matter for interoperability — is the subject of our companion guide to the IAB Audience Taxonomy.
Groups span consumer and business purchasing: Automotive (ownership, products and services), Travel and Tourism, Finance and Insurance, Software, Consumer Electronics, Real Estate, Education and Careers, Business and Industrial, and thirty more. Segments are where activation happens — not "Travel" but Cruise Travel, Hotels and Resorts, Travel Insurance; not "Finance" but Mortgage Lenders and Brokers, Retirement Planning, Credit Cards. The complete browsable structure, with every code and label, lives on the audience segmentation taxonomy page.
PI.auto_ownership Automotive Ownership
PI.travel Travel and Tourism
PI.finance_insurance Finance and Insurance
PI.software Software
PI.consumer_electronics Consumer Electronics
PI.real_estate Real Estate
PI.education_careers Education and Careers
PI.business_industrial Business and Industrial
PI.health_medical Health and Medical Services
PI.home_garden_services Home and Garden Services
PI.family_parenting Family and Parenting
PI.clothing_accessories Clothing and Accessories
+ 22 more groups · 283 segments total
The same PI.* signals feed very different workflows depending on who is holding them.
Buy the pages whose readership is in-market rather than the users a cookie once claimed were. A campaign for travel insurance runs across every page scoring Travel Insurance or Cruise Travel intent at medium-plus confidence — reaching in-market readers on 100% of traffic, including Safari and iOS where behavioral segments cannot follow.
Score accounts by what their buying committees actually read: enrich a target-account list with the intent profiles of the trade content, comparison pages and documentation their teams consume, and weight Software or Business and Industrial intent into your ABM prioritization — a content-side complement to surge-based intent vendors.
Curation platforms and SSPs package inventory into Deal IDs defined by intent: "auto in-market, high confidence, brand-safe" becomes a curated PMP assembled from the 102M-domain dataset plus per-URL scoring of fresh inventory — a sell-side data product with a derivation trail buyers can audit.
Publishers attach PI.* signals to their own inventory and pass them in the bid request as Seller Defined Audiences — IAB-aligned intent codes make the segments legible to any DSP that speaks the standard, turning editorial content about mortgages or EVs into packaged in-market supply.
Sales and partnerships teams use intent-scored domain lists as prospecting filters: every domain whose audience shows Logistics and Delivery intent is a lead list for a freight platform; every site with Retirement Planning intent is one for a wealth-management tool selling ad placements or integrations.
Intent tells you what an audience is shopping for; personas tell you who they are. Combining PI.* segments with our deterministic 1,667-persona layer — covered in the companion guide to audience personas in advertising — yields plans like "First-Time Homebuyer personas on pages with Mortgage Lenders and Brokers intent".
Intent data has a long history of overclaiming. These are the constraints of the content-inferred approach, stated plainly — because segments you can defend to a buyer, a lawyer or an auditor are worth more than segments you cannot.
A high-confidence New Vehicles signal on a page means its readership skews toward car shoppers — not that any specific visitor is one. Some readers of an EV comparison are journalists, competitors or the merely curious. For media buying this is the right resolution (media is bought in aggregate impressions), but it cannot replace user-level data for tasks that genuinely require it, like suppressing recent purchasers.
Inferred intent carries a low / medium / high confidence band, not a decimal probability — because the underlying evidence (how strongly content selects an in-market readership) does not support four significant figures. A mortgage calculator is high-band; a general personal-finance column is low-band or carries no intent at all. Thresholds are yours to set per use case: strict for guaranteed deals, permissive for prospecting.
No intent product — behavioral or contextual — predicts that a purchase will occur, and content-inferred intent does not claim to. It identifies where in-market attention concentrates. Conversion depends on the offer, the creative, the price and the moment; intent data improves the odds of being present when consideration happens, nothing more.
Thin, ambiguous or misleading pages yield weak or absent signals — by design, the system returns nothing rather than guessing. And because signals derive from content, a page that changes topic changes profile: per-URL scoring handles this in real time, while domain-level profiles refresh on the dataset's update cycle rather than instantly.
A single API call against an electric-SUV comparison review on an automotive publisher. Coded PI.* values on the left; the same segments rendered as labels for planning and packaging on the right.
POST /api/audience/segment.php
{ "query": "https://autoreviews.example/2026-electric-suv-comparison" }
// response — purchase_intent family (other families omitted)
{
"purchase_intent": [
{ "code": "PI.auto_ownership.new_vehicles",
"label": "New Vehicles",
"confidence": "high" },
{ "code": "PI.finance_insurance.insurance",
"label": "Insurance",
"confidence": "medium" },
{ "code": "PI.auto_products.automotive_parts_and_accessories",
"label": "Automotive Parts and Accessories",
"confidence": "low" }
],
"vocab_version": "1.0"
}
The review compares purchase prices and trims (high-band New Vehicles), discusses EV insurance costs in one section (medium-band Insurance), and mentions accessories in passing (low-band). Nothing else is asserted — no demographic guess is smuggled in as intent, and the same response carries demographics, life stage and personas as separate families with their own confidence bands. An auto OEM buys the high band; an insurer tests the medium band; an accessories retailer probably passes.
Paste a page into the live demo and watch the PI.* branch light up — in-market segments with confidence bands, alongside demographics, life stage, personas and B2B signals. No signup, no cookies, no tracking.