Enter a URL and get the likely age brackets, gender skew, income band, education level and life stage of its readership — inferred from the page's content, with no cookies, panels or user tracking involved.
Analyze what the site publishes. Pages about index-fund expense ratios, toddler sleep regressions, or enterprise Kubernetes carry strong signals about who is likely reading them. A content-inferred engine codes those signals into standardized attributes with banded confidence.
Paste a URL into the live demo, let the engine read the content, and get a structured demographic profile in seconds.
One application of cookieless audience segmentation — building audience intelligence from content rather than tracking people.
The alternative routes — publisher media kits, survey panels, tracking-based analytics — each answer the question partially, slowly, or only for the largest sites. We compare their trade-offs below.
Illustrative output: coded attributes with banded confidence scores.
Website demographics traditionally come from three places. Each is legitimate and each has structural limits worth understanding before you rely on it.
Media kits describe the audience a publisher wants to sell — whole-site level, updated infrequently, with no independent verification. Smaller sites often have no media kit at all.
Panels recruit users and project demographics onto sites. The method works for the head of the web but thins out fast — mid-tail and long-tail domains get tiny sample sizes or no coverage.
Cookie- and ID-based audience data requires observing individuals across sites. Safari, Firefox and iOS already block third-party cookies, making roughly 40%+ of traffic invisible.
| Approach | Coverage | Granularity | Freshness | Privacy exposure | Best for |
|---|---|---|---|---|---|
| Publisher media kit | Sites that publish one | Whole site | Updated yearly at best | None | Direct deals with large publishers |
| Survey / metered panel | Head of the web; thin in the tail | Site-level averages | Monthly-ish, lagging | Panelists consent; low | Benchmarking major sites |
| Cookie / ID analytics | Chrome-heavy subset of traffic | User-level where observable | Near real time | High; consent-dependent | Retargeting where IDs persist |
| Content-inferred (this tool) | Any URL or domain, 102M dataset | Per page or per domain | On demand, per request | None — no users observed | Planning, curation, competitive analysis |
These categories are complements. Content inference is the approach that works for every site — including competitors, long-tail placements, and niche blogs no panel covers.
Every attribute comes from a controlled, versioned vocabulary (v1.0) aligned with IAB Audience Taxonomy 1.1. Demographics are model-inferred with banded confidence; personas are deterministic. Full code lists are on the taxonomy page.
8 standardized brackets, e.g. 25_34. A page can score on more than one bracket.
5-point scale from strong male to strong female skew — a skew of readership, not a claim about individuals.
6 bands describing the likely household income range of the typical reader.
7 levels, from secondary through postgraduate, inferred from reading level and subject matter.
14 stages — student, young professional, new parent, empty nester, retiree and more.
Household composition and home-ownership signals — renter-leaning vs. owner-leaning readerships.
Employment type and, for B2B content, firmographics in LinkedIn-standard bands (role, seniority, size).
Urban, suburban or rural lean of the likely readership, where content signals support it.
Reads a single page at request time. Best when a domain hosts many audiences — e.g. a newspaper whose sports, finance and parenting sections reach different people.
Pre-computed profiles for scale work: scoring placement lists, enriching a CRM, or curating inventory across thousands of sites without issuing live requests.
Because age_bracket: 25_34 means the same thing in every response, data flows into a DSP, spreadsheet or clean room without translation. The vocabulary is versioned, so profiles stay interpretable over time.
The same query also returns interest groups (29 / 285 sub-interests), purchase intent (34 / 283 segments), and audience personas (1,667). Full surface on the audience segmentation page.
No account, no tag on the target site, no waiting for data to accumulate. The engine reads the page at request time and infers the audience from what it finds.
Go to the audience intelligence demo dashboard. It runs the same engine as the production API.
Paste any publicly reachable URL. A specific article gives page-level demographics; a homepage gives a domain-flavored view. For batch analysis, use the 102M-domain dataset.
It extracts main text, classifies against IAB categories, then infers readership from topic, vocabulary, reading level and dozens of other signals. No cookies, no visitor observed.
Results come as controlled-vocabulary codes: age brackets, gender skew, income, education, life stage — each with a low / medium / high confidence band.
The same response includes interest groups, purchase-intent segments and mapped personas. Together they form a complete audience brief for the page.
Planning media for readers of independent personal-finance publishers? Run a representative article — a guide to maximizing employer 401(k) matching — through the demo.
Interest values use INT.* codes and intent values use PI.* codes, rendered here as labels.
High-confidence matches: The 401(k) topic presupposes salaried employment with benefits — hence high confidence on 25–44, degree-educated, employed and middle-to-upper income.
Low-confidence signals: Home ownership and urbanicity are flagged low because the text carries only weak signals. The confidence bands make this transparent.
Actionable conclusion: This placement matches a 25–44, employed, mid-to-upper-income brief — and intent shows readers actively researching retirement products.
Paste a competitor's page, a placement from your last campaign report, or your own site — and see the coded demographic profile it returns.
A demographic profile is a planning primitive. These are the four workflows where teams use it daily.
Verify candidate placements actually reach the demographic in the brief — per page, not per publisher average. Score a full URL list against your target profile.
Profile a competitor's site or their ad placements to see whose readership they court. Any public URL is analyzable — no tag, no partnership required.
Publishers package inventory into demographic segments for the bidstream via the Seller-Defined Audiences framework — giving even untagged pages a sellable signal.
Enrich a CRM record with the demographic and firmographic profile of the company domain behind it. The 102M-domain dataset makes this a join, not a crawl.
Content-inferred demographics are powerful because they make a modest claim. Using them well means knowing the boundaries.
The profile describes the audience a page's content most plausibly attracts. It never identifies or stores data about any individual — that is what makes it privacy-safe by construction.
A 2,000-word specialist article gives the model far more to work with than a thin landing page. Low-band attributes on signal-poor pages should be treated as hypotheses.
A general-interest portal's finance and celebrity sections reach different readerships. Where a domain is heterogeneous, prefer per-URL lookups on representative pages.
Topic and reading level constrain age, education and employment well; they constrain home ownership or urbanicity only when the content addresses them directly.
For certified measurement of who actually visited a site last month, use panels and analytics. For a fast, consistent, privacy-safe demographic read on any page or domain — including the 99% of the web no panel covers — content inference is the tool built for the job.
Enter the site's URL into a content-inferred audience tool such as our live demo. The engine reads the page, classifies it against IAB categories, and returns age brackets, gender skew, income band, education, life stage and related attributes — each with a banded confidence score. No tag on the target site is required.
Every attribute carries a low, medium or high confidence band rather than a single point estimate. Attributes tightly constrained by topic and reading level — age, education, employment — typically come back high-band on substantive pages. Weakly signaled attributes like urbanicity are flagged lower so you can weight them accordingly.
Yes. Content inference derives the profile entirely from what the page publishes — no cookies, device IDs, fingerprinting or visitor observation. Safari, Firefox and iOS already block third-party cookies, leaving roughly 40%+ of traffic invisible to tracking-based methods. Content-based analysis works identically across all browsers.
Our engine returns 8 age brackets, a 5-point gender-skew scale, 6 income bands, 7 education levels and 14 life stages, plus household composition, employment, home ownership and urbanicity. All from controlled, versioned vocabularies aligned with IAB Audience Taxonomy 1.1. The same lookup also returns interests, purchase-intent segments and personas.
Audience measurement (panels, census analytics) reports who actually visited a site, but only for large enough sites and only at site level. Content-inferred demographics describe the audience a page is written for — available for any URL instantly, at page-level granularity, without observing any individual.
Yes — any publicly reachable URL can be analyzed, including sites you have no relationship with. Profile a rival's key pages, compare their inferred readership against your own, and see which demographic and intent segments they are publishing for.
Paste a URL into the live demo and get age, gender skew, income, education, life stage, interests, intent and personas back in seconds. No signup needed.