Context hint strategy: building targeting you can actually read

Tarun Kapoor, founder of Context Hints, seated at a wooden desk with a soft city light behind him.Tarun Kapoor Updated 1 September 2026 Full reference · 26,000 words

Writing context hints is easy. A language model will produce a plausible set in seconds, and our own generator will do it for nothing. What is hard is that the platform reports nothing back at hint level — no search-term report, no per-hint performance, no way to know which hint fired. That single constraint changes what the work actually is: the ad group becomes the unit of measurement, and the job stops being writing phrases and becomes designing a structure that can tell you something.

The short version

  • What it is: a research and architecture engagement — we harvest the language your buyers actually use, cluster it into coherent intents, and build the ad-group structure that makes those clusters readable.
  • The constraint that shapes everything: there is no search-term report and no per-hint performance data. You cannot see which hint fired. So the ad group, not the hint, is the unit of measurement — and it has to be built that way from the start.
  • What we are not selling: hint generation. A language model will write you plausible hints for nothing, and tools exist that do it with evidence attached. Use them. What they cannot do is read an outcome the platform does not report.
  • What it costs: from $3,999, one-off, fixed before we start.
  • The honest part: the full method is published below, and the context hints guide covers the theory in more depth than this page does. Both are free.

What is context hint research?

Context hint research is a one-off engagement that sources the language your actual customers use to describe the problem you solve, clusters that language into an ad group architecture carrying one intent per group, and writes the five to fifteen context hints each of those ad groups needs — the plain-language descriptions OpenAI matches against live conversations at ad group level in Ads Manager Beta. OpenAI’s own definition is that advertisers “provide context hints that describe the conversations, topics, or keywords where their products or services may be relevant”, and that these hints “help guide ad matching, but they are not exact-match keywords and do not guarantee delivery in specific conversations”. The research and the architecture are the engagement; the hints themselves are the smallest and cheapest part of the deliverable. Ours is from $3,999 as a one-off. The canonical free explainer of the mechanism sits in our guide to context hints, and this page assumes you have read it or will.

That definition puts the weight in an unusual place on purpose. Almost everything sold under the phrase “context hint help” is the last step — producing sentences — and producing sentences is the step that has already been automated. What has not been automated is knowing which sentences to produce, how many ad groups they should be split across, and how you will read the result once the money is spent. Those three questions are what this engagement answers.

What does “guide matching but not guarantee delivery” actually mean?

A context hint is a bid for consideration, not a booking. You are describing a class of conversation you would like to be considered relevant to; OpenAI decides, conversation by conversation, whether any ad is shown and which one wins. There is no mechanism to reserve a topic, no way to name a conversation you want to appear in, and no confirmation afterwards that you appeared in one. Advertisers do not get access to chats, chat history, memories or personal details — reporting is aggregated by design, and that is a privacy decision rather than a roadmap gap.

Delivery is also gated by conditions your hints cannot override. Ads serve to Free and Go plan users only — never Plus, Pro, Business, Enterprise or Education. Users declared or predicted to be under 18 are excluded. Conversations about politics, health and mental health carry no ads at all. A hint that beautifully describes a patient researching a treatment pathway will never deliver, however well written it is, and no amount of iteration will surface that fact. Knowing which parts of your demand are structurally unreachable before you write a hint set is worth more than a better sentence.

How is research different from writing a list of hints?

Writing a list of hints is a text-generation task and software does it competently. We publish a free one at the context hint generator, and paid tools that do the same thing for a few dollars a month exist. If what you need is fifteen plausible sentences from a product description, use one of them and keep your $3,999. That is a sincere recommendation, not a rhetorical setup.

What a generator cannot do is read your call recordings, your support tickets, your win/loss notes, your churn interviews and your sales objections, because it has never seen them. It writes from your product description, which is precisely why generated hints describe the product rather than the moment the buyer is in — the single most expensive and least visible failure in this channel, because the account looks fully configured while it is happening. It also cannot decide how many ad groups you should have, which is the decision that determines whether you can learn anything from the spend at all. The mechanics of hint construction are free in our guide to writing context hints. The judgement about what to write and how to split it is what you are buying here.

Is this the same as ongoing management?

No — research ends with an architecture and a hint set you own outright; management is what rewrites them every week afterwards. This engagement is a one-off from $3,999 and it terminates in a document, a build sheet and a handover. Our White-Glove ChatGPT Ads retainer is $1,499/month flat, covering up to $10,000 in managed monthly ad spend, with Tarun personally operating the account. If you have an in-house operator with genuine spare capacity, buying the research and keeping the operating is the correct and common outcome. If you have neither an operator nor a research problem, buy neither.

What it isWhat it producesWhen it is the right buy
Hint generation — software, free to a few dollars a monthA list of plausible sentences derived from your product description, in minutesYou have one product, one ad group, and a small enough budget that being roughly right is enough. Start with our free context hint generator.
Context hint research and architecture — this engagement, from $3,999 one-offCustomer language sourced from your own recordings, tickets and win/loss notes; an ad group map with one intent per group; hint sets sized to be readable; a measurement plan for a channel that gives no per-hint feedbackYou are spending enough that the cost of an unreadable account exceeds the fee, and nobody in the business has ever had to design an experiment they could not instrument.
Ongoing hint managementWhite-Glove retainer, $1,499/monthThe same rewriting done every week against live impression and conversion data, alongside bids, budgets, creative and measurementThe architecture exists and is right, and the problem is that nobody has time to operate it on the Mondays nobody is watching.
None of the aboveAn afternoon and a checklistMedia budget under roughly $1,500/month. Any fee dominates the media at that level. Use the free context hints guide and keep the cash in the auction.

The constraint that defines this discipline: no per-hint feedback

ChatGPT Ads has no search-term report and no per-hint performance data, so it is not possible to see which context hint caused an impression, which conversation your ad appeared in, or which hint produced a conversion. Ads Manager Beta has four sidebar sections — Campaigns, Tools, Billing, Settings — and no reports tab. There is no placement or “where ads showed” report, no auction insights and no ad-level A/B testing primitive. Every metric the platform gives you — impressions, clicks, spend, CTR, average CPC, average CPM, conversions, and the optional VTA (1d) column where eligible — is reported at campaign, ad group and ad level. None of it is reported at hint level, and no export, column setting or segment produces it.

This is the first thing to say on a page selling hint work, because it invalidates most of the way that work is normally described. Every consequence below follows from it.

What is the unit of measurement if not the hint?

The ad group is the unit of measurement, and you test hint sets rather than hints. A hint set is the whole multi-line block typed into one ad group; the platform reports what that block did in aggregate, mixed with whatever the ads in that group did. Change three hints inside a group of eight and the CTR that comes back describes the new set of eight, not the three you changed. Attribution to an individual hint is not available and cannot be reconstructed after the fact.

Two surfaces get mistaken for the missing report. Edit columns, under the three-dot menu, adds metric columns including individual conversion events. Segment controls Device and Country breakdowns only and will not add conversion-event columns. Between them you can slice an ad group by device and country; below that, the platform stops.

What partially replaces the missing granularity is downstream: the four URL macros — {campaign_id}, {ad_group_id}, {ad_id} and {ad_account_id} — substituted at delivery time and set through the three-dot menu rather than the creation flow. Tagged properly, your analytics and CRM carry ad-group-level outcome data Ads Manager will not surface; tagged at the wrong level, you get one undifferentiated blob of ChatGPT traffic. The precedence order and escaping rules are in our guide to ChatGPT Ads URL parameters.

Why does ad group architecture become the primary instrument?

Ad group architecture is the primary instrument because the ad group is the smallest object the platform reports on, which makes it the only object you can design an experiment around. In Google Ads, structure is a convenience: you can be sloppy about ad groups and recover the truth from the search-term report afterwards. Here there is no afterwards. The architecture you build before launch is the resolution at which you will be permitted to see your own results.

That inverts the order of the work: structure stops being what you tidy up after the hints are written and becomes what the hints are written to fit. The practical test is to write down, before launch, the sentence you want to be able to say in six weeks. If it is “buyers comparing a shortlist convert and problem-aware browsers do not”, then comparing-a-shortlist and problem-aware have to be separate ad groups, or that sentence stays unprovable no matter how long the campaign runs. Ad group boundaries that do not correspond to the questions you want answered produce a number and no knowledge.

The ad group is the smallest object the platform will tell you the truth about, so the ad group is the only place you can put a question.

Why is one ad group, one intent a measurement requirement rather than a style preference?

One ad group, one intent is a measurement requirement because a mixed-intent ad group returns one blended number for two different buyers, and the blend is not decomposable. Suppose — illustratively — an ad group carries hints describing both someone who has just realised they have a problem and someone comparing two shortlisted vendors, and it returns a 1.2% CTR and four conversions. No operation in Ads Manager tells you whether those four came from the comparison hints at a good cost or the problem-aware hints at a terrible one. Split into two ad groups, the same spend answers the question by itself.

The rule has a second consequence that reinforces the first. Relevance is computed from four inputs — context hints, landing page, ad title and ad copy — so the copy has to be right for the moment the hints describe. An ad group straddling two stages cannot have copy right for both, and the mismatch becomes a relevance penalty you cannot see and cannot price. The four-stage architecture we build against is problem-aware, researching options, comparing a shortlist, ready to act, under the Audience — Intent — Topic framework set out in our guide to OpenAI Ads intent targeting.

Why does adding more hints reduce what you can learn?

Adding hints to an existing ad group destroys the comparison you were running, because the reported result belongs to the set and the set has changed. The working range is five to fifteen per ad group: launch with five and watch impressions for five to seven days before touching anything.

“Just add more hints” is the default reflex when an ad group under-delivers, and it does usually increase impressions. That is exactly why it is dangerous: it fixes the symptom you can see and worsens the one you cannot. Each addition widens the set, dilutes the intent, and adds another candidate explanation for every number that follows. Three additions in one week means the following week’s CTR has at least four possible causes and no way to separate them. The fifteen-hint ceiling bites on diagnosis rather than delivery — a fifteen-hint ad group is not a targeting problem, it is an unreadable one.

The disciplined alternative is to change one thing per ad group per test window, hold the ads constant across the ad groups being compared, and run comparisons in parallel rather than sequentially, because there is no ad-level A/B testing primitive to do it for you. Parallel ad groups on equal bids and comparable budgets are the closest thing to a controlled test the platform supports.

What do you do about waste when there are no match types and no negatives?

There are no match types and no negative-keyword list in ChatGPT Ads. Nothing corresponds to exact, phrase or broad match, and there is no list of terms to exclude into. Negative hints are expected later; today there are two workarounds, and they are not equivalent.

A threshold rules the second route out for most advertisers who reach for it: a list needs at least 25,000 matched users, and OpenAI recommends 100,000 or more. Audiences also cannot be edited after creation, so a mis-specified exclusion is a new upload rather than a correction. Below that scale the tighter hint is not the preferred workaround, it is the only one. Our guide to custom audiences covers the file format and hashing rules.

What should you ask anyone promising per-hint optimization?

Ask them to show you where the data comes from. There is no per-hint reporting surface in Ads Manager Beta, no per-hint field in the CSV export, no per-hint dimension in the API, and no per-hint value returned by any of the four URL macros. A vendor offering to “optimize your highest-performing hints” or “pause the underperformers” is describing an operation the platform does not expose. The charitable explanation is that they are inferring hint performance from ad-group performance and have not said so; the uncharitable one is that they imported a Google Ads vocabulary wholesale and have not noticed the report is missing.

Two questions settle it. Which screen is that number on? And if two hints in the same ad group both fired, how does your process tell them apart? The correct answer to the second is structural rather than reported — you split them into two ad groups and read them separately. Anyone who produces a per-hint performance table should be asked to export it in front of you.

What you would check in a Google Ads search-term reportWhat is available in ChatGPT Ads insteadWhat that changes about the method
Which query triggered the impressionNothing. No search-term report, no placement report, no auction insights, no access to conversationsThe question has to be designed into the structure beforehand. Which ad group spent is the only proxy for which conversation you reached.
Cost by search term, sorted descending, to find spend that bought nothingSpend, clicks, CTR and conversions at campaign, ad group and ad level; cost per conversion derived as spend ÷ conversionsWaste is diagnosable only down to the ad group. Ad groups therefore have to be narrow enough that a wasteful one is identifiable as a whole.
Add negative keywords for the queries that wasted budgetNo negative-hint list. A tighter hint, or a campaign-level custom-audience exclusion needing 25,000+ matched usersExclusion is a rewrite, not an addition. Most accounts are below the audience threshold, so the sentence is the only lever.
Diagnose match-type creep — broad match pulling in adjacent categoriesNo match types at all. Breadth is a property of how the hint is writtenThere is no setting to tighten. Precision is authored, and it is authored once per ad group rather than tuned per term.
Spot query-to-landing-page mismatch from the terms reportRelevance computed from four inputs — context hints, landing page, ad title, ad copyYou audit the four inputs against each other for agreement instead of auditing queries against pages. The mismatch is inferred, never displayed.
Harvest new keywords from queries you did not think ofNo discovery surface. The hint field is a bare text box with no suggestions and no volume dataNew language has to be sourced outside the platform — from customers, calls, tickets and reviews. There is no way to discover it from spend.
Segment a term by device, then by hour, then by geographySegment controls Device and Country breakdowns only; Edit columns adds metrics, not dimensionsTwo breakdowns exist and both sit under the ad group. Everything finer is pushed into your own analytics via {ad_group_id}.
Attribute a conversion back to the keyword that earned itoppref credits a conversion to an ad interaction; URL macros credit a session to an ad groupoppref is how OpenAI credits a conversion to an ad; macros are how you credit a customer to an ad group.” Neither reaches the hint.

One further reason not to over-read what the ad group does tell you: AdVenture Media reports that roughly 40% of ChatGPT-ad-attributable conversions happen in the immediate post-click session and roughly 60% arrive later by paths the platform cannot see. That is their observation, not our result, and it should temper how confidently you act on a small ad-group difference. Closing that gap is a measurement build, covered in our conversion tracking guide, not a hint problem.

A translucent blue glass funnel on a white field receiving many fine threads of light at its wide mouth and emitting a single combined beam below, with no way to trace the outgoing beam back to any individual incoming thread.
Every hint in an ad group feeds one reported result. The platform returns the combined beam — impressions, clicks and conversions for the ad group — and no operation traces any part of it back to an individual hint. The only way to separate two threads is to put them in two different funnels before the campaign is switched on.

What this means for your account

Your ad group map is your measurement plan, decided before launch rather than discovered afterwards. Every question you want answered in six weeks must correspond to an ad group boundary you drew today, because the platform will never report below that line. Build with five hints on one intent, hold the ads constant while you compare, wait five to seven days on impressions before judging anything, and treat any offer of per-hint performance data as a claim to check rather than a feature to buy.

Hints are not keywords, and the difference is operational

Keywords match strings and context hints match situations, which means the two are not different syntaxes for the same instruction — they are different objects with different failure modes, different pricing behavior and different tooling. A keyword is a token that either appears in a query or does not; matching is lexical, governed by match types, and a keyword’s reach is the set of phrasings a match type permits. A context hint is a described circumstance evaluated against the meaning of a whole conversation; matching is semantic, ungoverned by match types, and a hint’s reach is however many ways people can express the situation you described. The structural comparison is set out in full in our context hints vs keywords guide; what follows is what the difference does to the work.

What changes when you are describing a situation rather than a string?

Three things change, and each of them retires a habit that Google Ads rewarded.

How do context hints set the price you pay?

Context hints are one of four relevance inputs, and relevance is a multiplier on your bid, so hint quality is a price lever rather than only a targeting lever. OpenAI runs a relevance-weighted second-price auction: ads are ranked by bid multiplied by relevance, and the winner pays one increment above the second-place effective bid. Relevance is computed from context hints, the landing page, the ad title and the ad copy. The image is explicitly not an input.

OpenAI’s own worked comparison makes the size of the effect concrete: a $2.50 bid at 0.9 relevance beats a $4.00 bid at 0.4. The arithmetic is 2.25 against 1.60 — the lower bid wins by a comfortable margin, on a 60% smaller number. Everything that follows about writing hints follows from that line. A hint set that describes the buyer’s moment accurately, backed by a landing page and ad copy that answer that same moment, does not merely reach better conversations; it reaches them at a bid your competitor cannot beat by outspending you.

Relevance inputWhere it is setWhat it has to agree with
Context hintsAd group, in a multi-line text boxThe moment described must be one the landing page actually serves. A hint promising a comparison, pointed at a homepage, is an internal contradiction the auction prices.
Ad titleAdThe stage the hints describe. A title written for a ready-to-act buyer under problem-aware hints is a mismatch inside your own ad group.
Ad copyAdThe same stage, in the buyer’s vocabulary rather than your product’s. Length guidance is in our creative guide.
Landing pageAd, as the landing page URLThe specific situation, not the brand generally. This is the input most often left as a homepage and the one that most often drags a good hint set down.
The imageAdNothing. The image is explicitly not a relevance input, so it cannot rescue a weak hint set and cannot be blamed for one.

Why does importing a Google Ads keyword list produce bad hints?

A keyword list is a record of the words people typed, stripped of the situation that made them type it, and a hint made by pasting a keyword into a sentence inherits that loss. There is no bulk keyword import in Ads Manager, but the informal version happens constantly: an operator opens their top-spending keyword report, converts each line into a phrase, and lands with fifteen hints that name product categories nobody is having a conversation about. The table below is illustrative — the keywords are generic examples, not any client’s account — and the pattern it shows is the one we find most often.

Google Ads keywordThe naive hint made from itThe hint that describes the buyer’s moment
crm software“CRM software”“A founder whose deals live in a shared spreadsheet is asking how to stop losing follow-ups now that two people are updating the same sales list.”
best project management tool“Best project management tools for teams”“A ten-person agency is comparing project tools after missing a client deadline because the brief existed in three places and none of them agreed.”
payroll software small business“Payroll software for small businesses”“A first-time employer is working out what they are legally responsible for withholding before they run payroll for their first two hires.”
emergency plumber near me“Emergency plumber near me”“Someone with water coming through a ceiling is asking what to shut off first and whether anyone can get out to them tonight.”
accounting software pricing“Accounting software pricing and plans”“A bookkeeper is trying to work out what accounting software actually costs at 500 invoices a month, because the published tiers stop at 100.”
crm comparison“CRM comparison and reviews”“A revenue operations lead has shortlisted two CRMs and is weighing migration effort and per-seat cost for a 40-person team without hiring an implementation consultant.”
how to reduce churn“How to reduce customer churn”“A subscription founder whose renewals slipped two months running is trying to establish whether the cause is onboarding or pricing before buying any tool.”

Four failure patterns are visible in the middle column. The naive hints name a category rather than a circumstance, so they compete for every conversation that mentions the category and stand out in none. They carry no stage, so the ad group they belong to cannot have coherent copy and the relevance score suffers on three of the four inputs at once. They use your vocabulary rather than the buyer’s — nobody in a chat window says “project management tool” before they have said “we missed a deadline”. And they collapse distinct intents into one line: “CRM software” silently contains the founder who does not yet know they need one and the buyer who has a purchase order waiting, which is the mixed-intent problem from the previous section arriving through a different door.

What shape does a good context hint take?

The canonical shape is a described scenario: a sentence or two naming who is in the conversation, what they are trying to do right now, and the topic and constraint they are doing it under. It is not a keyword and it is not a persona, and the distinction between those last two is where most decent marketers go wrong.

The persona shape fails for a specific mechanical reason rather than a stylistic one: the platform is not matching people, it is matching conversations. There is no age or gender targeting, no retargeting audience and no cross-device identity, so a hint asserting firmographics is asserting something the matching layer has no way to act on. What it can act on is the situation the sentence describes. The Audience — Intent — Topic template that produces the third shape reliably, with worked variants by funnel stage, is free in our guide to writing context hints, and fifty vertical examples sit in the main context hints guide.

How many hints per ad group, and how long do you wait?

Five to fifteen context hints per ad group is the working range, and the sequence is to launch with five and watch impressions for five to seven days before changing anything. Five is enough to accumulate delivery on a single intent and few enough that the set stays interpretable. Fifteen is a ceiling rather than a target — an ad group at fifteen has usually stopped being one intent.

The wait is not patience for its own sake. Attributed conversions take 24 to 48 hours to appear, so anything read on day one is noise, and impressions on a narrow hint set accumulate slowly enough that two days of data can invert on the third. The most common self-inflicted wound is a rewrite on day two, which resets the comparison and leaves you with a week of data describing three different sets. If impressions are near zero after seven days, the cause is usually not the hints: work the non-delivery checklist first — verification and billing complete, campaign and ads active, dates including today, ads through review — and note that bidding much below $3 wins little or no delivery whatever the hints say.

What does the hint input field actually give you?

The context hint field is a bare multi-line text box with no autocomplete, no suggestions, no volume estimator and no competitor research. There is no keyword planner because there are no keywords. Nothing in the interface tells you whether a hint you are about to write describes a conversation that happens a thousand times a day or never happens at all, and nothing tells you whether another advertiser has written the same one.

Operators coming from paid search underestimate what that removes. No volume estimate means a hint set cannot be sized before it runs — the first signal is the impression count five to seven days after launch, which is why the launch-with-five discipline exists. No competitor visibility means you cannot see who else is describing your buyer’s moment or infer a bid landscape from it. No suggestion engine means coverage is limited to the situations you were able to think of, which is exactly what makes sourcing language from real customers the highest-leverage part of this engagement rather than a nice-to-have. A generator — ours is free at the context hint generator — gives you the shape of a hint in seconds. It cannot tell you which situations your buyers are actually in, and neither can the field you are typing into.

The research method, end to end

Context hint research runs as a seven-stage sequence: harvest recorded buyer language, strip your own company vocabulary out of it, cluster what is left by the situation the buyer was in, map each cluster to one of the four intent stages and screen out the ones ChatGPT Ads cannot reach, draft a hint set per surviving cluster, architect ad groups and campaigns around those clusters, and write down in advance how the account will be read. Every stage produces one named artifact, and no stage can begin before the stage above it has produced it. The order is the method. Most accounts that look like they have done hint research have in fact done stage five on its own — someone sat down and wrote fifteen sentences — and that is why the hints describe the product.

The sequence is deliberately front-loaded. Roughly two thirds of the working time is stages one to three, before a single hint is written. That is the opposite of how the work is usually scoped, and it is the reason this is a research engagement rather than a copywriting one. A generation tool starts at stage five with your website as its input. Our context hint generator does exactly that, free, and it is a reasonable way to produce a first draft. What it cannot do is stages one to four, because the raw material lives in your call recordings and your helpdesk, not on your marketing site.

  1. Stage 1 — Harvest the raw language. We collect recorded sales calls, demo transcripts, support tickets, chat logs, reviews of you and of your competitors, community threads, and win/loss notes. Nothing is paraphrased and nothing is summarized at this stage; verbatim sentences are pulled with the speaker role and the date attached, because a sentence loses most of its value once it has been tidied. The artifact is a phrase corpus — a flat list of what people actually said, each line traceable back to its source. Two rules govern the harvest: only the buyer’s words go in, never the seller’s, and a sentence is captured with enough surrounding context to tell what the person was trying to do when they said it.
  2. Stage 2 — Strip the company vocabulary. Every corpus arrives contaminated, because buyers who have already spoken to your sales team have been taught your words. We mark and remove terms that only exist inside your company or your category: product names, internal module names, the acronym your industry uses for itself, the noun you invented for the problem. Where a buyer used your term, we look for what they said before they were given it — usually earlier in the same call. The artifact is a cleaned phrase bank, with a second short list of the terms that were removed. That removed list is useful on its own: it is a map of the vocabulary your ads should not be written in.
  3. Stage 3 — Cluster by the buyer’s situation. The cleaned phrases are grouped by the circumstance being described rather than by the words used to describe it. This is manual work done by reading, not by keyword frequency, because the whole point is that two people describing the same situation frequently share no vocabulary at all. Each cluster gets a plain-English name written in the buyer’s register, a count of how many distinct sources it drew from, and the three or four verbatim quotes that define it. The artifact is a clustered situation map. A typical engagement produces six to twelve durable clusters from a healthy corpus, of which four to eight are worth an ad group.
  4. Stage 4 — Map to intent stage, then screen for reachability. Each cluster is assigned to exactly one of the four stages — problem-aware, researching options, comparing a shortlist, ready to act — on the evidence of what the buyer was doing when they spoke, not on where you would like them to be. A cluster that will not sit in one stage is two clusters. Then the reachability screen, which most hint work skips entirely: ChatGPT Ads serve to Free and Go plan users only, never to Plus, Pro, Business, Enterprise or Education; users declared or predicted to be under 18 are excluded; and conversations about politics, health and mental health carry no ads. A cluster that lives inside any of those is struck here rather than discovered as silence eight weeks into a campaign. The artifact is a stage-tagged cluster list with exclusions marked and the reason recorded.
  5. Stage 5 — Draft the hint set for each cluster. Only now does anyone write a hint. Each surviving cluster gets five to fifteen hints, which are paraphrases of the same situation rather than five different situations, because five to fifteen per ad group is the working range and the variants exist to cover natural language variation. Each hint is built against the Audience — Intent — Topic framework and checked against the verbatim quotes that define its cluster. We launch with five and hold the rest in reserve, since under-delivery is usually a coverage problem before it is a bid problem. The artifact is a hint set per cluster, plus the reserve.
  6. Stage 6 — Architect the ad groups and campaigns. One cluster becomes one ad group. Campaign membership is decided by what is actually set at campaign level — objective, conversion event, budget, dates, countries, platforms, and custom-audience inclusions and exclusions — because clusters that need different values for any of those cannot share a campaign. Each ad group gets its bid strategy chosen deliberately rather than accepted by default, its ads written to answer the same situation the hints describe, and a landing page that a person in that situation would recognize as the right page. The artifact is a build sheet: campaigns, ad groups, hints, ads, landing pages and bid strategies, in the shape a bulk upload will accept.
  7. Stage 7 — Define the read. Before anything launches, we write down what evidence will be examined, at what cadence, and what decision each outcome triggers. This matters more here than in a keyword channel because Ads Manager reports at campaign, ad group and ad level only — impressions, clicks, spend, CTR, average CPC, average CPM and conversions — and does not break performance out by individual hint. There is no search-term report and no placement report. The ad group is therefore the smallest unit of evidence you will ever have, which is what makes the architecture in stage six a measurement decision. The artifact is a one-page read plan naming the metric, the threshold, the waiting period and the action for each ad group.
StageArtifact it producesHow you know it is finished
1. HarvestVerbatim phrase corpus, sourced and datedTen consecutive new sources produce no phrase you have not already captured
2. StripCleaned phrase bank, plus the removed-terms listNo line in the bank contains a word that exists only inside your company
3. ClusterNamed situation clusters with source counts and defining quotesEvery phrase sits in exactly one cluster and no two clusters fail the merge test
4. Map and screenStage-tagged clusters, unreachable ones struck with reasonsEach cluster sits in exactly one of the four stages and has passed the reachability screen
5. DraftFive to fifteen hints per cluster, with a launch fiveEvery hint traces to a verbatim quote in its own cluster
6. ArchitectBuild sheet: campaigns, ad groups, hints, ads, pages, bidsNo ad group carries two intents and no campaign carries two settings requirements
7. Define the readOne-page read plan with thresholds and actionsEvery ad group has a named metric, a waiting period and a decision rule

What exactly is a cluster?

A cluster is a set of phrases describing the same buyer situation — the same trigger, the same constraint and the same decision being made — regardless of the words used. It is not a topic, not a keyword group and not a persona. A persona says who someone is; a cluster says what just happened to them and what they are now trying to settle. “Facilities managers” is a persona. “A flat roof that has leaked in two consecutive storms and a decision to make between another patch and a full replacement before winter” is a cluster, and only the second one can be written into a hint that describes a conversation.

A finished cluster has three parts on the page: the trigger that started the conversation, the constraint that will disqualify most vendors, and the decision the buyer is trying to reach. If any of the three is missing, the cluster is not finished, and a hint written from it will come out generic in exactly the way that costs money — because relevance in OpenAI’s auction is computed from context hints, landing page, ad title and ad copy, and a vague hint gives the other three nothing to agree with.

How do you know when two clusters are actually one?

Two clusters are one cluster if the same ad, pointed at the same page, would be the best available answer to both. That is the whole test, and it has three parts, all of which must come back negative before the two stay separate.

Clusters that fail the test and merge are not wasted. Their phrases become hint variants inside the surviving ad group, which is precisely what the five-to-fifteen range exists to hold. The reverse case is more common than people expect: one cluster that is really two, revealed when the same sentence turns out to be said by two different roles with two different constraints, or by the same role at two different stages. “We are looking at replacing our supplier” from an owner who has already been let down twice and from a buyer running a routine annual review are the same words and different clusters, and they need different ads.

How much raw language do you need before clustering means anything?

Our working floor is around forty recorded conversations, or roughly 300 to 400 distinct verbatim phrases, before clustering produces anything you should spend money against. Below about fifteen conversations you get personas rather than clusters, because there is not enough repetition to tell a real pattern from one articulate customer. A situation has to be described by at least five unrelated sources before we will name it as a cluster; below that it is an anecdote, and anecdotes make narrow hints that never accumulate impressions.

The stopping rule is saturation, not a quota. We keep harvesting until ten consecutive new sources produce no phrase that does not already sit in an existing cluster. In a business with a narrow product and one buyer type, saturation can arrive at thirty conversations. In a business selling four products to three buyer types, it will not arrive at eighty, and the honest answer is to scope the research to the one or two segments that matter this quarter rather than to pretend the corpus is complete.

There is no volume data to check your clusters against

Context hints are entered in a bare multi-line text box at ad group level. There is no autocomplete, no suggestions, no volume estimator and no competitor research, because there is no keyword planner and there are no keywords. Nothing in Ads Manager will tell you how many conversations a cluster corresponds to. Cluster frequency in your own corpus is the only sizing signal available, and it is a proxy for demand rather than a measurement of it — it tells you what your buyers talk about, weighted by who happened to reach your sales team. Treat it as a ranking, not a forecast, and let the first five to seven days of impression data correct it.

What this means for your account

The value in this sequence sits in stages two, three and four, which are the stages no software performs and no marketing site contains. Hints written without them are your own vocabulary handed back to you in sentence form, and they will read as complete in the interface while matching conversations that were never going to convert. If you want the underlying mechanics before you buy anything, our guide to context hints covers what they are and OpenAI ads intent targeting covers how the matching behaves.

Where the language actually comes from

Buyer language for context hints comes from six sources, all of them inside your own business or in public places your buyers already write in: recorded sales calls and demo transcripts, support tickets and chat logs, reviews of you and of your competitors, community threads, win/loss notes, and your own search-query data from other channels. None of it comes from ChatGPT. Advertisers do not get access to chats, chat history, memories or personal details, and there is no search-term report and no placement report in Ads Manager, so the platform will never hand you the conversations your ads appeared in. Every word you write into a hint box has to be earned somewhere else first.

The six sources are not interchangeable. Each one is a recording of buyers at a particular moment in their relationship with you, which makes each one systematically biased in a direction worth naming before you use it. The failure that follows from mixing them carelessly is a hint set that describes your existing customers to strangers.

Why are recorded sales calls the highest-value source?

Recorded sales calls and demo transcripts are the highest-value source because they capture your buyer describing their own problem, in their own words, before you have reframed it for them. The first sixty to ninety seconds of a discovery call is usually the single most useful stretch of language your company owns: the buyer has not yet adopted your terminology, they are explaining why they took the meeting, and they almost always name the trigger event and the constraint in the same breath. That is a cluster arriving fully formed.

What it biases toward is people who already agreed to a call. That is a filtered population — further along than a stranger in a chat window, pre-qualified by whatever your team screens for, and shaped by the questions your reps ask. Later in the call the contamination gets worse, because by minute twenty the buyer is repeating your words back to you. Harvest the opening, not the demo.

To harvest: pull recordings from whatever your team already records, confirm the consent and retention position with whoever owns it, transcribe, redact names and identifying details, then read the openings rather than the summaries. Automatic call summaries are worse than useless here, because summarizing is the exact operation that destroys the vocabulary you came for. This is the slowest input in the engagement and the one worth the time.

What are support tickets and chat logs good for?

Support tickets and chat logs are post-purchase language, which makes them strong for retention and expansion angles and misleading for acquisition. The vocabulary in a ticket assumes your product exists, assumes the writer already knows what your features are called, and describes a problem inside your world rather than the problem that made them buy. Used for an acquisition hint, that produces a sentence only a customer could have written — and customers are not who the ad has to match.

There is one exception worth mining: the ticket or chat that describes what the person was trying to do when they signed up, which surfaces most often in onboarding conversations and in cancellation threads. Cancellation reasons in particular tend to restate the original job in plain language, because the writer is explaining what they thought they were buying. Sentiment also skews — tickets exist because something went wrong, so the corpus over-represents friction.

To harvest: export from the helpdesk with dates and the ticket type, filter to onboarding, pre-sales and cancellation categories, and read those rather than the technical queue. Effort is low compared with calls, since the text already exists.

What do reviews tell you that a sales call cannot?

Reviews are the words people use when nobody is selling to them, which makes them the cleanest available record of decision criteria stated in the buyer’s own frame. Your competitors’ reviews are usually more useful than your own, because a complaint about a competitor is a specification for what your buyer wanted and did not get, written by someone with no reason to be diplomatic. Read the three-star reviews first: five-star reviews are gratitude and one-star reviews are anger, but three-star reviews are trade-offs, and trade-offs are constraints.

What reviews bias toward is the extremes of experience and the structure of the review prompt. A platform that asks “what do you like most?” produces answers shaped like feature lists. They are also retrospective: a reviewer is reconstructing a decision they made months ago, so the trigger event is often lost or tidied into something more rational than it was.

To harvest: read rather than scrape, take verbatim sentences with the star rating attached, and keep the ones that describe the situation before purchase rather than the experience after it. Effort is low and the yield per hour is high, which is why this is the first source we open when a business has no recordings.

How useful are community threads and forums?

Community threads capture the question as it was first asked, before any vendor was involved, which makes them the best available source for problem-aware language. Someone posting “has anyone dealt with this” in a trade forum is doing in public what a large share of ChatGPT users are doing in private, and the sentence structure is close to what a person types into a chat window. The replies matter as much as the post, because that is where the constraints get argued out.

The bias is heavy and predictable: forums over-represent technical, self-serve and price-sensitive buyers, over-represent whoever posts most, and reflect the demographics of the platform rather than of your market. A category whose real buyers are fifty-year-old operations managers will look completely different on a developer forum than it does on a trade association board. Match the forum to the buyer, and discount anything that only ever appears in one thread.

To harvest: search your category the way a buyer would phrase the problem, not the way you would name the product, and read threads end to end. Effort is medium, and the noise ratio is the highest of the six sources.

What do win/loss notes and recurring objections add?

Win/loss notes are the best source for the comparing-a-shortlist and ready-to-act stages, because they record which alternatives were actually on the list and what tipped the decision. Objections that recur across multiple deals are constraints in disguise: an objection raised in one deal is a personality, and the same objection raised in nine is a cluster boundary. Losses are more informative than wins, for the obvious reason that a loss names the thing you were missing.

The bias is that these notes are written by the seller, after the fact, in the seller’s vocabulary and with the seller’s explanation attached. Losses are also chronically under-documented, because nobody enjoys writing them up, so the surviving record skews toward wins and toward the reasons that reflect well on the process. Treat the note as a pointer to a conversation worth finding in the call recordings rather than as language you can lift.

To harvest: pull the closed-lost reasons from the CRM for the last two to four quarters, count the recurrences, and use the top handful to go and find the corresponding call recordings. Effort is medium, and most of it is in reconciling free-text fields that were filled in inconsistently.

Can you convert search-query data from other channels?

Search-query data from your other channels is useful for checking coverage and dangerous as source material, because it is keyword language rather than conversational language. A query is a compressed, telegraphic string typed at a search box by someone who has learned to talk to a machine. A ChatGPT conversation is a sentence typed by someone talking about their circumstances. Converting the first into the second naively — taking your top search terms and pasting them into the hint box — produces exactly the failure everything above is designed to prevent: hints that read as keyword fragments, describe categories rather than situations, and give the auction nothing specific to match against.

Used properly, this data does two jobs. It confirms that a cluster you found in the calls also shows up in demand elsewhere, and it surfaces vocabulary you had not heard — a term buyers use that never came up in a call because your reps steer past it. Neither job involves copying a query into a hint. Every term that survives has to be rebuilt into a situation first, which means going back to the corpus to find someone describing that term as a circumstance.

To harvest: export the query report from whichever channel you already run, sort by conversions rather than volume, and use the top terms as a checklist against your clusters. Effort is low; the conversion step is where the work actually is. The structural difference between the two models is unpacked in our guide to context hints.

SourceWhat it is good forWhat it biases towardEffort to harvest
Recorded sales calls and demo transcripts The buyer’s unreframed description of the problem; trigger events and constraints stated together People who already agreed to a call; your rep’s question set; your vocabulary after the first few minutes High — access, consent check, transcription, redaction, then reading openings rather than summaries
Support tickets and chat logs Retention and expansion angles; the words for things going wrong Post-purchase framing; existing-customer assumptions; friction and negative sentiment Low — exportable text, but needs filtering to onboarding, pre-sales and cancellation categories
Reviews, yours and competitors’ Decision criteria stated when nobody is selling; competitor gaps as specifications The extremes of experience; the review platform’s prompt structure; retrospective tidying Low — read rather than scrape; three-star reviews first
Community threads and forums Problem-aware language; the question as first asked, before any vendor contact Technical and self-serve buyers; the loudest posters; the platform’s demographic Medium — high noise, manual reading, threads end to end
Win/loss notes and recurring objections Shortlist and ready-to-act stages; what the real alternatives were The seller’s memory and framing; wins over losses; inconsistent free-text fields Medium — mostly reconciliation, then tracing objections back to recordings
Search-query data from other channels Coverage checking; surfacing vocabulary you have not heard Keyword syntax rather than sentences; machine-directed phrasing Low to export, high to convert — every term must be rebuilt as a situation

What if you have none of this?

A business with no recorded calls, no reviews and no support history has very little raw material, and we will say so rather than invent hints from your marketing site. This is the most common reason we decline the engagement. The marketing site is the one input guaranteed to be written in your vocabulary, by you, about your product — running it through stages one to four produces nothing, because there is nothing in it that did not come from you in the first place. Hints derived from it will describe your offer, match conversations about your category in general, and cost you money at a stage of the funnel where nobody is ready to act.

The honest sequence in that situation is to build the corpus before buying the research. Start recording discovery calls and keep the openings. Ask ten recent customers what they were trying to do the week before they found you, and write down the answer verbatim rather than summarized. Read fifty of your competitors’ reviews and keep the three-star ones. Three to six weeks of that is usually enough to make the engagement worth its fee, and it costs nothing but attention. We would rather tell you to wait than take $3,999 for a document built on the copy in your own footer.

There is one intermediate case. A business with no calls but a large public review presence, or an active community where its category is discussed daily, has a usable corpus of a different shape — weaker on triggers, stronger on constraints. We will take that engagement and say up front which of the four stages will be under-served by it, which in practice is almost always the ready-to-act stage.

What this means for your account

The quality ceiling on your context hints is set by the quality of your recordings, not by the skill of whoever writes the sentences. An account with two hundred transcribed discovery calls and an account with a homepage are not the same engagement and should not produce the same document. Before you buy hint research from anyone, ask what corpus they intend to build it from, and treat “we will review your website and your competitors” as the answer it is.

Sorting intent into ad groups

Context hints are organized into ad groups by intent stage, using four stages that describe how far into a decision the buyer has travelled: problem-aware, researching options, comparing a shortlist, and ready to act. The rule is one ad group, one intent, and it is enforced because the ad group is the smallest unit Ads Manager reports on, not because a tidy account is pleasant to look at. Language looks materially different at each of the four stages, and an ad group that mixes two of them cannot be read, cannot be bid correctly, and cannot be answered by a single ad.

How does buyer language change across the four stages?

The reliable tell is what the sentence contains. Problem-aware language contains a symptom and no category. Researching-options language contains a category and no vendor. Comparing-a-shortlist language contains vendors. Ready-to-act language contains a deadline, a budget, or a specification.

StageWhat the buyer is doingWhat a hint looks likeWhat the ad should offer
Problem-aware Describing a symptom, with no category name and no vendors in mind Names the trigger and the mess it created, in the buyer’s words, with no product noun An explanation of what is happening and why, on a page that does not ask for a demo above the fold
Researching options Learning how the category works and what it costs Names the category and the specific question being asked about it A straight answer to that question — how it works, what it typically costs, what it does not do
Comparing a shortlist Weighing two to four named alternatives against a constraint Names the constraint that decides it, and the situation that produced the constraint The comparison on the buyer’s terms, including where you are the wrong choice
Ready to act Settling timing, onboarding, contract and approval Names the deadline, the specification or the approval that has to be cleared The specific next step, with the friction removed and the timeline stated

Why is one ad group, one intent a measurement requirement?

One ad group, one intent is a measurement requirement because Ads Manager does not report performance by individual hint. Metrics are available at campaign, ad group and ad level — impressions, clicks, spend, CTR, average CPC, average CPM and conversions — and there is no search-term report, no placement report and no auction insights. The ad group is therefore the finest grain of evidence that will ever exist in this channel, and whatever you put inside one becomes a single undifferentiated number.

The consequence is concrete. Put a problem-aware cluster and a ready-to-act cluster in the same ad group and the reported CTR is a weighted average of two populations that behave nothing alike, the reported cost per conversion is an average of a cheap outcome and an expensive one, and any change you make afterwards moves a number you cannot attribute. You will not be able to tell whether the ad group improved because the ready-to-act half got better or because the problem-aware half stopped delivering. In Google Ads you would resolve that with a segment or a query export. Here, Edit columns adds metric columns and Segment controls device and country breakdowns only. Neither will separate two intents that you chose to merge.

The same logic disqualifies the other common shortcut, which is splitting ad groups by product line while leaving intent mixed inside each one. That produces ad groups that are readable as revenue lines and unreadable as decisions, and since the hint is the input the auction actually scores, it optimizes the wrong axis. Split by intent first. Split by product only where the products have genuinely different buyers.

How many ad groups should you actually run?

Run the smallest number of ad groups that keeps each intent separate, because every additional ad group divides the same conversion volume into a smaller sample and there is a floor below which no ad group can be read at all. The arithmetic below is illustrative — it is a worked example, not a projection for your account — but the shape of it holds everywhere.

Suppose an account spending $12,000 a month and settling at an average CPC of $4.00 — the middle of the $3–$5 starting bid range OpenAI suggests, and a plausible clearing price given that a bid cap is the maximum you can pay in an auction rather than what you will pay — with a landing page converting at 2.5%. That is 3,000 clicks and roughly 75 conversions a month across the whole account, which is about 17 a week before it is divided by anything.

Ad groupsClicks per ad group / monthConversions per ad group / monthConversions per ad group / weekWhat you can read
13,00075~17Account-level trend; no intent-level insight at all
21,500~38~8.7A slow directional read on each half, over several weeks
4750~19~4.3Monthly comparison only; weekly numbers are noise
6500~13~2.9A single conversion moves cost per conversion by a third
8375~9~2.2Nothing you can act on inside a quarter
12250~6~1.4Zero-conversion weeks are the norm and mean nothing

Illustrative arithmetic, evenly split. Two things make the real picture worse than the table. Budget is set at campaign level rather than ad group level, so an even split is not something you can allocate — the auction distributes within the campaign, and some ad groups will take far more than their share while others starve. And daily budgets have been a seven-day average since 27 July 2026, with maximum daily spend at twice the daily budget and maximum seven-day spend at seven times it, so daily numbers swing by design and only the weekly total is meaningful. Add the 24 to 48 hours it takes for attributed conversions to appear, and a Tuesday morning read of yesterday’s ad group performance is not a read at all.

The practical floor: below roughly 15 to 20 conversions a week, an ad group should stay on Manual: Max bid. Maximize results is automated bidding and it needs conversion volume to have anything to learn from; it shipped in August 2026 and is default-on for eligible new ad groups, which means leaving it alone is a decision you have made rather than a decision you have avoided. On the illustrative numbers above, every configuration except a single undivided ad group sits below that floor — which is the real answer to how many ad groups to run at $12,000 a month. Split for meaning, keep the count low, and accept manual bidding as the normal state rather than a failure.

What should the ad groups be called?

Name each ad group after the cluster it carries, in the buyer’s language, with the intent stage in the name. Ad group names must be unique within a campaign, and they are the only place in the account where the reason an ad group exists is recorded — there is no notes field, no label system and no description. An ad group called Fleet cameras tells the next person nothing. An ad group called Second collision — insurer pressure — problem-aware tells them the cluster, the trigger and the stage, which is enough to decide whether a new hint belongs in it.

The naming discipline has a second, harder use downstream. The four URL macros — {campaign_id}, {ad_group_id}, {ad_id} and {ad_account_id} — return IDs rather than names, so anything you tag into your own analytics arrives as a number. You have to keep a mapping table from ID to cluster name yourself, and it has to be maintained by whoever creates ad groups. Without it, the segmentation you built the whole architecture to produce lands in your analytics as an undifferentiated block of ChatGPT traffic, and the intent structure exists only inside Ads Manager where you cannot join it to revenue. Build the mapping table in the same sheet as the hint sets, at stage six, before anything is created.

When you genuinely need spend guaranteed per intent

Ad groups do not have budgets. If a particular intent must receive a defined amount of spend regardless of what the auction prefers, the only instrument is a separate campaign, because budget, objective, conversion event, dates, countries, platforms and custom-audience inclusions and exclusions are all campaign-level settings. That buys you control and costs you volume per unit, since you have now divided the same conversions across two campaigns as well as several ad groups. Do it when the intents have genuinely different economics, not to make a report look balanced.

Four translucent blue glass panels standing in a row as a gentle ascending stair on a white field, each taller and more saturated than the last, with pools of light beneath them growing brighter from left to right.
Problem-aware, researching, comparing, ready to act. Each stage has its own language, and each gets its own ad group so the read stays clean.

How a hint is actually written

A context hint is one sentence that describes the situation a buyer is in, written in the buyer’s vocabulary, containing the constraint or worry that makes the moment real, and covering exactly one situation. It is built against the Audience — Intent — Topic framework, which is a completeness check rather than a template: all three layers must be present, but the finished sentence has to read like something a person would say, not like three fields joined with commas. The full template library, with worked examples by funnel stage, is in our guide to writing context hints; what follows is the discipline that decides whether the sentence you produce is worth the fee.

What does Audience — Intent — Topic look like worked properly?

Audience is who is having the conversation, concretely enough to exclude people. Role, seniority, company size or stage, sector, and where relevant the circumstance that defines them. “Operations managers” excludes nobody. “Operations managers running twenty to eighty vans” excludes almost everybody, which is the point. The richness here is what lets the matching engine connect the hint to a conversation where the person has already stated similar facts about themselves.

Intent is what they are trying to do right now, and the verb carries it. Researching, working out, comparing, switching, replacing, choosing, shortlisting, approving. Each verb implies one of the four stages, which is how you check that a hint belongs in the ad group you put it in. If you cannot write an active verb because you do not know what the buyer is doing, you have a persona rather than a cluster and the problem is upstream in the research.

Topic is the category, the sub-category and the constraint, and the constraint is the part that does the work. The category is what a hundred competitors also sell. The constraint is what disqualifies most of them: the integration that must exist, the deadline that cannot move, the contract length that will not be signed, the skill the team does not have. A hint without a constraint describes a market. A hint with one describes a conversation.

Worked properly, those three collapse into a single sentence with the trigger in it. Take a cluster named “second at-fault collision, insurer asking questions”. Audience: operations managers running twenty to eighty vans. Intent: being asked to demonstrate what they are doing about driver behavior. Topic: vehicle cameras, with the constraint that it has to produce something an insurer will accept. Assembled: operations managers running twenty to eighty vans who have just had a second at-fault collision this year and are being asked by their insurer to show what they are doing about driver behavior. One sentence, three layers, one situation, no product name.

The four rules

Before and after: twelve rewrites

The following pairs are illustrative rewrites, not client work. In each case the left column is the kind of sentence that gets typed into the hint box on a Friday afternoon, and the right column is what the same business should have written after the research.

BusinessHint as writtenRewrittenWhat changed, and why
Commercial roofing contractor Commercial roofing services, flat roof repair, TPO installation, 25-year warranty Facilities managers of single-storey commercial buildings whose flat roof has leaked in two consecutive storms, deciding between another patch repair and a full replacement before winter Keyword fragments replaced with a situation and a decision. The repeat leak is the trigger and “before winter” is the constraint; the warranty is an offer detail nobody says out loud
Self-storage operator Affordable self-storage units near you, first month free, 24/7 access People moving out of a rented flat three weeks before the new place is ready, needing somewhere to put a one-bedroom of furniture without signing a twelve-month contract A promotion is not an intent. The rewrite names the gap between two tenancies, the volume, and the contract-length worry that actually decides which operator gets called
Fleet camera and telematics software AI-powered fleet safety platform with real-time driver coaching and video telematics Operations managers running twenty to eighty vans who have had a second at-fault collision this year and are being asked by their insurer to show what they are doing about driver behavior Product taxonomy out, trigger event in. Fleet size makes the audience concrete and the insurer supplies the pressure that makes the moment real
Speciality coffee roaster, wholesale Speciality coffee beans, single origin, direct trade, freshly roasted Independent cafe owners unhappy with the consistency of their current bean supplier who want to change roaster without retraining staff on a different espresso recipe The original describes a product a hundred roasters claim identically. The rewrite names the buyer as wholesale rather than retail, the switching trigger, and the switching cost that is blocking the decision
Independent e-bike retailer Best electric bikes 2026, free delivery, price match, finance available Commuters with a five to eight mile ride each way working out whether an e-bike would genuinely replace the car journey, unsure where they would charge and store it in a flat A paid-search headline pasted into a hint box becomes a substitution decision with its two real objections attached. Offer terms and the year removed — neither appears in how anyone describes their commute
Corporate travel management Corporate travel management, duty of care, TMC, savings up to 20% Office managers at companies of fifty to three hundred staff whose travel is booked on personal cards and consumer sites, told to bring expenses and traveller safety under one process Acronyms buyers never say are gone. The rewrite names the current state being replaced and the internal mandate that started the search; a savings percentage describes you, not the conversation
Uniform and workwear supplier Branded workwear and PPE supplier, embroidery, bulk discounts Someone opening a second site for a food business who has to kit out fifteen new staff in hygiene-compliant uniforms in one order that arrives before the opening date Bulk-discount language is a sales term. The rewrite supplies the expansion trigger, the headcount and the deadline — the three things that decide which supplier wins the order
Commercial mortgage broker Commercial mortgage broker, competitive rates, whole of market, fast decisions Owners of a small business whose commercial mortgage comes off its fixed rate in the next few months, wanting to understand what refinancing involves before their current lender makes an offer Every adjective in the original is one a competitor claims word for word. The rewrite names the timing event that creates the conversation and the specific question being asked at that moment
Powered access and plant hire Powered access hire, scissor lifts, boom lifts, cherry pickers, nationwide Site managers who have found mid-job that the ladder plan will not pass the safety check and need a scissor lift on site within two days with the operator ticket sorted An equipment list becomes the moment the equipment is needed — unplanned, urgent and compliance-driven, which is what separates a hire conversation from a purchase
Password manager for small teams Enterprise-grade password management with SSO, SCIM and zero-knowledge encryption The accidental IT person at a twenty-person company who has realised the whole team shares logins in a spreadsheet and that two people who left still know the passwords Feature acronyms are how you describe yourself. The rewrite gives the discovery moment and the specific fear that makes it urgent; “enterprise-grade” also contradicted the audience it was aimed at
Driving school Driving lessons, first lesson free, high pass rate, automatic and manual Adults in their thirties who never learned to drive, fitting lessons around a full-time job, nervous about being in a car with a stranger for the first time The original addresses nobody in particular. The rewrite names an under-served segment, the scheduling constraint, and the worry that decides whether they book at all; pass-rate claims belong in the ad
Donor management software for small charities Nonprofit donor management platform with automated journeys and Gift Aid reclaim A fundraiser at a charity of under ten staff running the donor list in a spreadsheet, who has just missed a renewal deadline because of it and needs something the volunteers can use too The rewrite names the organization size, the failure that triggered the search, and the volunteer-usability constraint that eliminates most of the market before price is discussed

How specific is too specific?

A hint is too specific when the situation it describes is rarer than your delivery needs, and too broad when the ad group stops being readable. Both failures are silent, and they have different symptoms, which is how you tell them apart.

Too narrow shows up as no delivery. Hints guide matching but are not exact-match targeting rules and do not guarantee delivery in specific conversations, so a hint that is too tight does not return an error or a warning — it simply produces very few impressions, or none, and the ad group sits there looking configured. Watch impressions over five to seven days after launch. If an ad group is under-pacing, the first move is coverage rather than bid: add the reserve hints from stage five, since under-delivery is usually a coverage problem before it is a bid problem. The other narrow-hint trap is the reachability screen — a cluster that reads perfectly but sits inside an audience the platform does not serve will look exactly like a too-narrow hint.

Too broad shows up as an unreadable ad group. The impressions arrive, the CTR looks unremarkable, the conversions are a blend, and no change you make produces a legible result. A hint broad enough to match several situations pulls several populations into one number, and because there is no per-hint reporting you cannot see which of them you bought. The second cost is on the auction side: relevance is scored from context hints, landing page, ad title and ad copy, and a hint that describes three situations cannot agree with an ad that addresses one of them.

The practical calibration is length and count. One sentence, roughly the length of something a person would actually say when explaining their circumstances to someone who asked. A hint that needs a comma-spliced list of three products is three hints. A hint that fits on half a line has almost certainly lost the constraint. And the count is its own diagnostic: if a cluster cannot produce five distinct hints without repeating itself, the cluster is too narrow to be an ad group; if it produces more than fifteen genuinely different ones, it is two clusters that have not been separated.

The box will accept anything you type

Context hints are entered at ad group level in a bare multi-line text box. There is no validation, no character warning, no duplicate detection, no suggestion, no volume estimator and no competitor research. Nothing in the interface will tell you that a hint describes your product rather than a conversation, that two of your hints are the same sentence in different words, or that one of them belongs in a different ad group. The discipline has to come from the process that produced the sentence, because the tool will not supply any of it. That is also why hints are reviewable in bulk: they are pipe- or semicolon-delimited in a single cell in the bulk upload format, so the whole set can be read side by side in a spreadsheet before it goes anywhere near Ads Manager.

Working without negative keywords

There is no negative-keyword list in ChatGPT Ads, and there are no match types. Negative hints are expected in future but they do not exist today, which means you cannot tell the platform which conversations to stay out of. You have two instruments for narrowing delivery: writing a tighter hint, and excluding a custom audience at campaign level. That is the complete set. Anything else described as an exclusion control is a misreading of a feature that does something different, and the most expensive of those misreadings is the bid multiplier.

InstrumentWhat it actually doesWhat it cannot doWhen to reach for it
A tighter hint Changes what the ad group is described as matching, at ad group level Guarantee that any specific conversation is excluded — hints guide matching, they are not rules Almost always. Irrelevant delivery is usually a description problem, and this is the only instrument that addresses the description
Custom-audience exclusion Removes a list of matched users from a campaign’s eligible population Exclude a topic, a situation or anyone you cannot supply as hashed identifiers in volume Excluding known people: existing customers, current trialists, churned accounts you do not want back
Splitting the ad group Separates two intents so each gets its own hints, ads and bid Reduce total irrelevant delivery on its own When the irrelevant half is really a second cluster you failed to separate in the research
Bid multiplier (0.1x–10x) Adjusts what you bid for users in a matched audience, at ad group level Exclude anybody, or include anybody. It never gates eligibility in either direction Weighting bids toward a known-valuable audience — never as an exclusion or a targeting control

What do you need to know about custom audiences before you build one?

A custom audience is created at Settings → Audiences and applied at campaign level as an inclusion or an exclusion. The mechanics that decide whether this instrument is available to you at all are the size thresholds: the minimum is 25,000 matched users, and OpenAI recommends 100,000 or more. Matched, not uploaded — a list of 30,000 email addresses does not become a 30,000-user audience. Most B2B advertisers cannot clear 25,000 matched users from any list they own, which means exclusion is simply not available to them and the hint is their only instrument. Establish that before you plan a strategy around it.

The file rules are strict and worth getting right the first time. CSV or TXT, up to 500 MB and up to 5,000,000 identifiers, with one identifier type per upload — emails and phone numbers cannot be mixed in a single file. The optional header must be exactly email, phone_number, email_sha256 or phone_number_sha256. Processing takes 20 to 30 minutes, and the audience moves through statuses of Ready, Processing, Upload pending, Indexing, Publishing, Too small or Failed on the way.

Two constraints change how you should plan. Audiences cannot be edited after creation — there is no adding to one, no removing from one and no refreshing one in place, so every update is a new audience that has to be created, processed and re-applied to the campaigns that used the old one. Name them with a date from the first upload, keep the source file, and expect to rebuild on a schedule rather than to maintain. And the same audience cannot be both included and excluded in one campaign, so a plan that reads “target our customer list but exclude the churned ones” needs two separately built audiences, not one used twice.

How do you prepare the identifiers?

Normalize first, then hash. That order is not negotiable and getting it backwards silently destroys your match rate, which then presents as an audience that is too small rather than as an error.

The two errors we see most often are hashing an already-hashed value, which happens when a warehouse export is pre-hashed and someone hashes it again on the way out, and uppercase hex output, which several standard libraries produce by default. Both give you a file that uploads cleanly, processes normally and matches almost nobody. The full walkthrough is in our guide to ChatGPT Ads custom audiences.

Why is a bid multiplier not an exclusion tool?

Bid multipliers range from 0.1x to 10x and never gate eligibility. A 10x multiplier includes nobody who was not already eligible, and a 0.1x multiplier excludes nobody who was. Where several multipliers match the same user, the highest one wins and they do not stack. All a multiplier does is change what you bid for people who were going to be reachable either way.

Treating 0.1x as a soft exclusion is a common and expensive error. The reasoning goes that bidding a tenth as much for an audience will effectively remove it, and it does not: those users remain fully eligible, the ads still serve to them whenever the reduced bid clears the auction, and because the auction is relevance-weighted rather than a straight price race, a low bid with high relevance still wins impressions. What you have actually built is a segment you are bidding badly for while continuing to buy. The second error is the mirror image — setting 10x on a valuable audience and treating that as targeting, then wondering why delivery did not shift toward it. If you need someone excluded, exclude them at campaign level. If you need someone included and nobody else, that is a hint problem, not an audience problem.

What do you do when neither instrument fits?

When delivery is irrelevant and neither a tighter hint nor an audience exclusion resolves it, work through four moves in order before concluding the channel is at fault.

  1. Rewrite the hint against the cluster, not against the symptom. Irrelevant delivery is nearly always a description problem: the hint describes a situation more broadly than the one you meant, usually because the constraint is missing. Add the constraint and the trigger back in before touching anything else.
  2. Split the ad group. If the irrelevant traffic is coherent — the same wrong kind of buyer arriving repeatedly — it is a second cluster that your research merged. Separating it gives you a readable ad group for the traffic you want and a decision to make about the other, which may be worth keeping at a different bid on a different page.
  3. Fix the other three relevance inputs. Relevance is computed from context hints, landing page, ad title and ad copy, and the image is explicitly not an input. A landing page written for a broader audience than your hint will pull delivery broader than your hint, regardless of how tight the sentence is. Make all four agree before you conclude the hint has failed.
  4. Price it in, or stop. Some irrelevant delivery is structural and cannot be removed with the instruments that exist today. At that point the question is arithmetic rather than architecture: does the ad group still clear your break-even CPC once the waste is included? Break-even CPC is your target cost per acquisition multiplied by your landing page conversion rate, so a $50 target at a 4% conversion rate is a $2.00 break-even. If it clears, keep it and re-examine when negative hints ship. If it does not, pause the ad group. Pausing a losing ad group is a legitimate answer, and it is often the only honest one.

What none of this can do is what a negative keyword list does. There is no retargeting audience, no cross-device identity, and no age or gender targeting, so the exclusions imported from a Meta or Google playbook mostly have no equivalent here. Plan the account on the assumption that precision comes from the hint and from the architecture around it, and treat exclusion as a narrow tool for known people rather than as a way to steer topics.

What this means for your account

Because subtraction is barely available, the hint has to be right at launch in a way that a keyword list never has to be. In a keyword channel you can start broad and carve the waste away over ninety days using the query report. Here there is no query report to carve with, and the only real instrument is the sentence you wrote before you spent anything. That is the entire argument for doing the research first, and it is why this engagement is priced as research rather than as copy.

A broad soft field of blue light on a white surface meeting an upright glass plate with a single narrow slot, and only one tight sharp band of light passing through to the other side.
With no negative-keyword list, the instrument is the aperture itself: a tighter hint admits less, and that is the only real lever you have.

How to test hint sets when you cannot see hints

In ChatGPT Ads the unit of test is the ad group, never the individual hint, because Ads Manager reports impressions, clicks, spend, CTR, avg CPC, avg CPM and conversions at campaign, ad group and ad level, and reports nothing at all per hint. Every question of the form “did this hint work?” is unanswerable in this platform and always has been. Every question of the form “did this ad group’s hint set work?” is answerable — but only if the ad group was built so that it could be read, and most were not.

That single constraint decides the whole methodology. A hint set is not a list of independently measurable bets. It is one object, five to fifteen lines long, that either describes a coherent conversation or does not, and the ad group is the meter attached to it. Design the account so the meter reads something, or accept that you are guessing with a budget attached.

What makes an ad group a readable unit?

An ad group is readable when you can say in one sentence what it is for, and when no other ad group in the account would answer that sentence the same way. Three properties get you there, and all three are structural decisions made before any traffic arrives.

Name the ad group after the intent rather than the product, because the name is the only human-readable label the CSV export carries. The four URL macros — {campaign_id}, {ad_group_id}, {ad_id} and {ad_account_id} — return IDs and not names, so keep a mapping table from day one or your analytics will show you a result you cannot attach to a hypothesis.

What counts as changing one variable?

There is no ad-level A/B testing primitive in ChatGPT Ads, and no experiment primitive of any kind. A test here is a structural arrangement plus a disciplined read, not a feature you switch on. Nothing splits traffic for you, nothing holds a cell constant, and nothing warns you when the two arms stopped being comparable. Everything that must be held equal, you hold equal by hand.

Across two arms of a hint test, the following have to match: the bid strategy and the max bid, both set at ad group level; the ads themselves, since ad title and ad copy are two of the four inputs the auction reads for relevance, alongside context hints and the landing page; and every campaign-level setting, because objective, conversion event, budget, start and end dates, countries, platforms and custom-audience inclusions and exclusions all live on the campaign. The image is the one asset you can vary freely — the image is explicitly not a relevance input.

Then there is the confound the platform hands you for free. Budget is a campaign-level control and is not divided evenly between the ad groups underneath it. Two ad groups in one campaign compete for the same money, so the arm that wins delivery accumulates data faster and the arm that loses looks weak when it was mostly starved. Putting the arms in one campaign controls every setting and surrenders control of spend; putting them in separate campaigns controls spend and multiplies the settings you must keep identical by hand. When the answer matters, use separate campaigns with equal daily budgets and check every campaign field twice. When you only want to know which language gets served at all, one campaign is fine.

How long does a hint test have to run?

A hint test needs at least two whole seven-day blocks before it is worth reading, and four weeks before a conversion comparison is worth defending. Two platform mechanics set that floor, and neither is negotiable.

Allow 24 to 48 hours for attributed conversions to appear. Day-one data is not a weak signal, it is noise, and it is noise shaped exactly like a broken pixel. Discard the trailing 48 hours of any window before you read it, every time, including the window you are excited about.

Since 27 July 2026 the daily budget is a seven-day average: maximum daily spend is 2× the daily budget and maximum seven-day spend is 7×. OpenAI’s own example is seven days at $100/day totalling $700, spending $140 on Tuesday and $60 on Wednesday. A single expensive day is therefore a pacing artifact rather than evidence of anything, and any read shorter than seven whole days is a read of the pacing algorithm. Cut your windows in whole seven-day blocks on the same weekday boundary, in the account’s own time zone — time zone is set once at account creation and cannot be changed, and it is what draws the day lines in every report you will use.

The first read, at five to seven days, is a delivery read and nothing more. It answers “is this ad group eligible and being served”, which is a genuine question with a real failure mode behind it. It does not answer “is this hint set better than that one”.

Which confounds are specific to this platform?

Four confounds are peculiar to ChatGPT Ads and each of them has quietly voided tests we have been shown.

How many conversions do you need before a difference means anything?

The arithmetic below is illustrative, built from round numbers and a rough rule of thumb, and your own cost per click and landing-page conversion rate are the only ones that decide your answer. It is here to show the shape of the problem, not to give you a threshold to quote.

Illustrative arithmetic. Suppose two ad groups each spend $2,000 over thirty days. At $4 per click — inside OpenAI’s own $3–$5 starting guidance — that buys 500 clicks per arm. At a 2% landing-page conversion rate, each arm produces about 10 conversions. Now suppose one arm returns 10 and the other returns 14. That reads as a 40% improvement, and somebody will put it on a slide.

Illustrative arithmetic, continued. A serviceable rule of thumb for counts of this kind is that the noise on a count is roughly its square root: about ±3 on ten, about ±4 on fourteen, and roughly ±5 on the difference between them. A gap of four inside a noise band of five is not a finding. To make a 30% difference stand out you need counts where the gap outruns the noise — 100 against 130 is a gap of 30 against noise of about 15, which is finally worth arguing about. At the same $4 click and 2% conversion rate, 100 conversions per arm is roughly $20,000 of media per arm.

That is why our pilot sizing exists: a real pilot at $10,000–$30,000, sized to produce roughly 50 to 100 conversions, is the point at which testing becomes an activity rather than a ritual. For context on the variance you are fighting before you start, First Page Sage’s 12 June 2026 analysis put ChatGPT Ads conversion rates between 0.2% and 5.8% across 19 industries, and published CTR figures for the channel disagree wildly — 0.91%, 3.8–4% and 1.5–6% all appear in the trade press. Track your own; the benchmarks cannot settle anything for you.

At small budgets, most hint tests are noise

An account spending $2,000 to $3,000 a month cannot run a hint test that would survive contact with a statistician, and no methodology fixes that. Below roughly $1,500/month in media, running the account yourself is usually the right call, and the job is not testing — it is getting one coherent, correctly measured account live.

What to do instead: fewer, larger structural bets, read over longer windows. Rather than four ad groups differing by a phrase, run two that differ in kind — problem-aware language against ready-to-act language. Rather than reading weekly, read at four weeks and take a directional verdict. Rather than testing hints, test the architecture, because the architecture is the decision you cannot cheaply reverse.

Which test designs are actually available?

Six arrangements are worth running. The minimum-run column is our threshold rather than an OpenAI specification, and every one of these is something you construct and analyze yourself, because the platform provides no experiment mechanism to run any of them.

Test designWhat it can tell youWhat it cannot tell youMinimum run
Two ad groups, one campaign, hint sets differing by buying stage, identical ads and bid strategy Which stage of the buying process your language reaches at all, and which one clicks. The cheapest useful arrangement. Anything about relative efficiency, because the shared campaign budget is not split evenly and the losing arm is starved rather than beaten. Two whole seven-day blocks for a delivery and CTR read.
Two ad groups, separate campaigns, equal daily budgets, identical ads, bid strategy and campaign settings A genuine comparison of two hint sets on equal money. The strongest design available. Nothing about individual hints. Also nothing reliable if any campaign field drifted apart, and there is no alert when one does. Four weeks before comparing conversions; longer if the arithmetic above says your counts are too small.
Sequential rewrite of one ad group’s hint set, before and after A directional read when you have only one ad group’s worth of budget. Useful for a wholesale rewrite of an obviously bad set. Causation. Time is the control, so seasonality, competitor entry and every platform change in the window are inside your result. Four weeks before, four weeks after, with no other change in either window.
Structural split — one broad ad group divided into two narrower ones Whether the broad group was hiding two different intents. Often the highest-value move in the account. Which of the two new groups would have won on the old budget, since splitting halves the data rate in each arm. Four weeks, and expect to need longer than the ad group you split.
Incremental edit — adding or removing one or two hints in a live ad group Very little on its own. Its honest use is maintenance: retiring lines that never accumulated impressions. Attribution of any change in performance to the edit. The set is the unit; you changed the set. Two whole seven-day blocks, and only if nothing else in the ad group moved.
Hint count at the edges — a five-hint set against a fifteen-hint set on the same intent Whether your ad group is impression-starved because the set is too tight, or unfocused because it is too wide. The optimal count. There is no such number, and the working range is five to fifteen for a reason. Two whole seven-day blocks for the impression read; four weeks for anything downstream.

What this means for your account

Before you run a hint test, write down the two arms, the one thing that differs, the date you will read it, and the result that would change your mind. If you cannot fill in the last of those, you are not testing, you are collecting numbers to justify a decision already made. And if the illustrative arithmetic above says your arms will produce ten conversions each, do not run the test — make the bigger structural bet instead and read it in a month.

Reading the outcome

You read the outcome of a hint set from seven ad-group metrics and one supplemental column, because that is the entirety of what Ads Manager exposes: impressions, clicks, spend, CTR, avg CPC, avg CPM and conversions, plus VTA (1d) where the account is eligible. There is no per-hint breakdown, no search-term report, no placement report and no auction insights. What you are doing is inference from a small number of aggregate figures, and doing it honestly means knowing exactly which question each figure can answer.

Cost per conversion is not a first-class metric in all views — derive it as spend ÷ conversions. Do the division yourself, in the export, and write the formula down so that everyone reading the sheet is dividing the same two numbers. Three views are available: the table, the Insights charts and the CSV export. Three export modes exist, and the third is the one most operators miss: current table view, cumulative values, and Export daily values for a genuine day-by-day file. That last one is a different export rather than a filter, and it is what you need to cut clean seven-day blocks.

What does each metric actually say about your hints?

Why does view-through stay out of the read?

View-through conversions are supplemental evidence and must never enter a hint-set verdict, because the platform itself keeps them out of every number you would use to judge one. VTA (1d) is counted when someone converts within one day of an eligible impression and no qualifying click receives credit for the same conversion; where both qualify, the click wins. Four rules govern it, and all four exist to stop you double-counting.

  1. It is not in your Conversions total. The column stands alone. Anyone summing the two columns to make a month look better has destroyed the account’s baseline, and rebuilding one takes longer than the month they were flattering.
  2. It does not touch performance math. No effect on cost per acquisition, no effect on conversion rate, no effect on bidding or billing.
  3. oCPC does not optimize toward it. The Conversions objective bills per valid click and optimizes toward clicks likelier to drive your conversion goal — not toward impressions that precede one.
  4. The window is fixed at one day, is not configurable, and is independent of the click window. Lengthening how long you count clicks does not lengthen how long you count views.

Find it in the attribution breakdown on hover over Conversions, as an optional standalone column via Edit columns, and in CSV where available. Read it as a sanity check on whether your hints are surfacing at all in a category with a long consideration path. Never put it in the numerator of anything.

What does each signal pattern most likely mean?

The table below is how we triage an ad group from limited signal. It is a ranking of likelihood, not a diagnosis — check the cheap confound in the third column before you rewrite anything, because more hint sets have been rewritten to fix a bid problem than the other way round.

Signal patternMost likely hint problemWhat to change
High impressions, low CTR The hints select a real conversation and the ad does not answer it. The hint set is probably describing a topic rather than a buyer in a moment. Rewrite the ad title and copy to answer the conversation the hints selected, before touching the hints. If CTR stays flat, the hints are describing an audience that was never going to click.
Low impressions, at any CTR The hint set is too narrow, describes conversations that rarely happen, or sits below the five-to-fifteen working range. Check the bid first — bidding much below $3 wins little or no delivery, and a bid problem here looks exactly like a hint problem. Then broaden one dimension of the set, not all of them.
High CTR, no conversions The hints select curiosity rather than intent, or select a stage upstream of the one the landing page is built for. Confirm measurement before you conclude anything: the conversion column must be added via Edit columns, the event must match exactly, and you must have waited 24 to 48 hours. Then move the hint set one stage later, or move the page one stage earlier.
Rising avg CPC at an unchanged max bid Relevance is falling relative to whoever else is bidding on the same conversations. Your hints describe ground somebody now describes better. Rewrite hints, title, copy and landing page to agree with each other. Three of those four are the relevance inputs you control directly; the image is not an input.
Conversions arriving, CPA above break-even The intent is correct and the economics are not. This is a pricing finding, not a targeting one. Work the arithmetic: break-even CPC equals target CPA multiplied by landing-page conversion rate, and maximum allowable CPA is roughly average order value × gross margin × close rate. Then decide whether to cut the bid or retire the intent.
Two ad groups with near-identical numbers The hint sets overlap. You are reading one audience twice and calling it a comparison. Rewrite one set so the two are mutually exclusive in language, not merely differently named. Until then, no test between them can produce a result.
One ad group taking nearly all the campaign’s spend Possibly nothing about hints at all. Campaign budget is not distributed evenly between ad groups. Separate the ad groups into their own campaigns with their own daily budgets before concluding the starved one is worse.

Edit columns or Segment — which one adds what you need?

Edit columns and Segment are different controls and they are not interchangeable. Edit columns, reached through the three-dot menu, adds metric columns including individual conversion events under Conversions & events and the optional VTA (1d) column. Segment controls Device and Country breakdowns only and will not add conversion-event columns. A meaningful share of “our conversions are not showing” reports are somebody looking for a conversion event in the Segment menu, finding Device and Country, and concluding the platform does not report it.

The Device breakdown carries a second trap worth stating plainly. Targeting offers iOS App, Android App and Web, where Web covers both desktop and mobile web; reporting groups device into Mobile and Desktop only, with mobile web counted under Mobile. You cannot read your platform selections back out of the device breakdown, so if platform targeting differs between two ad groups, that difference is invisible in every report you have.

How do you reconcile ad clicks with analytics sessions?

Ad clicks and analytics sessions will not match, and they are not supposed to. Clicks measure ad interactions. Sessions depend on page load, redirect chains, consent handling, browser blocking, UTM handling, attribution windows and time zones, and every one of those subtracts. A gap is expected; an unexplained gap that changes month to month is a finding.

Reconcile properly: same date range, same time zone, and check campaign-level and ad-level activity in the CSV rather than eyeballing the summary. Remember that the account’s time zone was fixed at creation and cannot be changed, so if your analytics property runs on a different zone your day boundaries have never lined up and never will — compare whole weeks, not days. Carry ad-group identity across with the four URL macros, put them in values rather than keys, do not URL-encode the braces, and keep the ID-to-name mapping table, because the macros return IDs.

Decide which number is your system of record before you look at either of them. Choosing afterwards is not reconciliation, it is picking the one that flatters the month.

Two further facts belong in any honest read. Cross-tool numbers will not agree and are not supposed to, which is a reason to nominate a system of record rather than to keep hunting for the tool that agrees with you. And the structural attribution gap is real: AdVenture Media’s analysis puts roughly 40% of ChatGPT-ad-attributable conversions in the immediate post-click session, with roughly 60% arriving later by paths nothing in the stack can trace. A hint set judged only on platform-reported conversions is being judged on the minority of its output. The economics of reading that properly are worked through in our guide to measuring ChatGPT Ads ROI, and the campaign-side mechanics in our guide to ChatGPT Ads conversion campaigns.

Cadence: how often hints should change

Context hints should be read weekly, edited weekly at most, and rewritten wholesale about once a month. Weekly is the rhythm this channel rewards because the seven-day budget average makes a week the smallest honest unit of comparison, and because in a channel roughly four months old — measured from the 5 May 2026 self-serve launch — the accumulated reading of your own ad groups is the only proprietary knowledge available. Nobody has a five-year data set. The advertiser who has looked at their ad-group table fifty times knows things the advertiser who has looked at it four times does not.

Weekly is also, precisely, the thing that stops happening. A recurring task with no deadline and no named owner loses every week to a task that has both, and the hint review has neither by default. It produces no artifact anyone chases, it has no due date, and skipping it costs nothing visible this week. We have opened accounts where the hint sets carry the launch date and the account carries six months of spend, and in every one of them the owner knew it needed doing.

What should change weekly, monthly, and never reactively?

Three tiers, and the discipline is in respecting the third one.

The third tier has a cost that compounds in a way most operators never quantify. Every reactive change restarts the clock on a clean seven-day block, and a block interrupted in the middle is not a shorter read, it is no read at all. An account touched twice a week does not accumulate fifty-two weekly readings over a year — it accumulates none, because no window in it was ever left alone long enough to mean something. That is how an account can be worked on constantly, by a competent person, and still know nothing about itself after six months. Restraint is the productive part of the cadence, and it is the part that feels like doing nothing.

Why should you never change hints and bids in the same week?

Changing hints and bids in the same week destroys your ability to attribute either result, and it does so twice over. The obvious way is that two causes now compete to explain one change in cost per conversion, and nothing in the platform can separate them. The less obvious way matters more: a bid change alters which auctions you win, so it changes the mix of conversations your new hint set is being judged on. You have not run one confounded test, you have changed the population your test was sampling from, halfway through.

The working rule is one lever per week, alternated. Bid changes move in 15–25% increments and budget changes in 20–30% increments, so neither is a small nudge that can be slipped in alongside something else. If the week’s change is a hint change, the bid does not move; if the week’s change is a bid change, the hints do not move, and the following week’s read belongs to the bid.

How should hint sets drift as an account matures?

Hint sets should start broad and tighten as ad-group reads accumulate, because at launch you are learning what language this market actually uses and later you are exploiting what you learned. There is no keyword planner to consult first — no autocomplete, no suggestions, no volume estimator and no competitor research — because there are no keywords. The account is your research instrument, so it has to be pointed widely enough to return something.

  1. Launch broad, at five hints, and watch impressions for five to seven days. The question is eligibility, not efficiency. An ad group that never accumulates impressions has told you nothing except that the conversation is rare or the bid is under the delivery floor.
  2. Weeks two to six: widen where nothing served, retire where nothing coincided with delivery. Move toward the middle of the five-to-fifteen range with lines that describe one conversation from several angles.
  3. Months two and three: split rather than broaden. By now the reads are accumulating, and the highest-value move is almost always taking an ad group that has quietly acquired two intents and making it two ad groups with two sets of ad copy.
  4. Month four onward: tighten deliberately. Precision costs you reach, and you should only pay that price with evidence in hand. Tightening an ad group that has produced fifty conversions is optimization; tightening one that has produced four is guessing with extra steps.

The shape of that drift is worth stating as a direction rather than a schedule, because the schedule depends on your spend. A mature account typically has more ad groups than it launched with and no wider a hint set in any of them — growth in the account shows up as more meters, not as bigger ones. If your hint sets are getting longer over time while your ad group count stays flat, the account is drifting toward failure pattern four rather than maturing, and the monthly review is the place to catch it.

Why does a hint set written in June no longer describe September?

A hint set decays for three independent reasons, and none of them announce themselves in Ads Manager. The first is that language moves: the phrasing your buyers used to describe a problem in June is not what they use once a category has a new name, a new default tool, or a new controversy attached to it. The second is competitive: the auction is relevance-weighted second price, so when a competitor enters describing the same conversations more precisely, your clearing price rises and your delivery falls without you having changed a single line.

The third is that the platform itself moved five times in the four months after self-serve launched. Daily budgets became seven-day averages on 27 July 2026. Maximize results shipped in August 2026, default-on for eligible new ad groups. Platform targeting shipped in the same month, as did VTA (1d) in reporting and usable URL macros. Automatic Advanced Matching was enabled on existing Web pixels on 17 August 2026. A hint set written in June was written against a different set of defaults, a different budget-pacing model and a different conversion count, and nothing in the interface flags that the assumptions underneath it have all changed.

What this means for your account

Decide who owns the weekly read before you buy anything, including this. If nobody will own it, a one-off strategy engagement still leaves you better off than an unarchitected account — but it will decay, and you should price that in rather than be surprised by it in month five. If you want the weekly read to be someone else’s standing obligation, that is what our White-Glove ChatGPT Ads retainer is, at $1,499/month covering up to $10,000 in managed monthly ad spend. We would rather tell you which of the two you need than sell you the more expensive one.

The failure patterns we find most often

Seven hint-set failure patterns account for most of the damage we find in live ChatGPT Ads accounts, and six of the seven are invisible in the interface because the account looks fully configured in every one of them. There is no warning state for a bad hint set. The box accepts whatever you type, the ads serve or do not, and the numbers are consistent with several different explanations at once. What follows is the symptom you can see, the cause underneath it, and the fix — in the order we usually find them.

1. Hints written as keywords

Symptom. The hint box contains comma-separated fragments: crm software, crm pricing, best crm for small business. Impressions are erratic, CTR is unremarkable, and the ad group behaves as though targeting were only loosely connected to it.

Cause. The set was written by somebody with a decade of paid search behind them, importing a mental model that has no counterpart here. OpenAI is explicit that hints “guide matching but aren’t exact-match targeting rules”, and there are no match types, no negative-keyword list, no bulk keyword import and no delivery guarantee. A two-word fragment is not a precise instruction in this system — it is a very short, very ambiguous sentence, and ambiguity is the one thing a matching model punishes.

Fix. Rewrite every line as something a person would actually say while describing their situation, structured against Audience — Intent — Topic. Length is not the point; specificity about the person and the moment is. Our guide to writing context hints carries the before-and-after examples, and what context hints are covers the underlying model.

2. Hints that describe the product rather than the buyer’s moment

Symptom. The most expensive pattern on this list, and the hardest to see. Impressions accumulate, CTR is acceptable, spend is orderly, and conversions arrive at a rate that is disappointing without being alarming. Nothing in the account looks broken.

Cause. The hint set was written from the product page. It describes features, categories and the vocabulary of the company rather than the situation of the person — and people describe their problems long before they describe your solution. Because the auction reads context hints, landing page, ad title and ad copy together, a product-shaped hint set usually sits alongside product-shaped copy and a product-shaped page, so all four inputs agree with each other and are all pointed at the wrong end of the conversation.

Fix. Rewrite from the buyer’s side, using their words rather than yours. The raw material is already in your business: recorded sales calls, support tickets, the first two lines of inbound enquiries, and lost-deal notes. Pull thirty of them and take the phrasing verbatim. If the resulting hints sound less polished than your marketing copy, that is usually evidence you got it right.

3. Hints so narrow that nothing serves

Symptom. Five to seven days after launch, the ad group has a handful of impressions and no read at all. The account is spending — just not here.

Cause. Either the set describes a conversation that is genuinely real but genuinely rare, or it sits below the five-to-fifteen working range, or the bid is under the delivery floor and the hints are innocent. That last case matters: bidding much below $3 wins little or no delivery, OpenAI’s own starting guidance is $3–$5, and low impressions caused by a $1.20 max bid look identical to low impressions caused by an over-tight hint set.

Fix. Check the bid before touching the language. Then broaden along exactly one dimension — usually the topic, keeping the audience and intent fixed — and give it another two seven-day blocks. Broadening everything at once converts a narrow ad group into failure pattern four, which is worse, because a starved ad group is at least readable as starved.

4. Hints so broad that the ad groups overlap

Symptom. Three or four ad groups returning suspiciously similar CTR and avg CPC, with one of them taking most of the campaign’s spend. Every comparison you attempt comes out inconclusive, and it has done so for months.

Cause. The sets overlap, so the ad groups are drawing from one pool of conversations. Because budget sits at campaign level and is not divided evenly, whichever ad group the system finds most efficient absorbs the money, and the others look weak when they were simply outbid by their own account. Nothing here is a targeting result; it is an accounting artifact.

Fix. Make the sets mutually exclusive in language, not just differently named. The cleanest partition is by buying stage — problem-aware, researching options, comparing a shortlist, ready to act — because those four states use visibly different vocabulary and because each one justifies different ad copy. If two ad groups still cannot be told apart after the rewrite, they are one ad group and should be merged.

5. Hint sets that have never been rewritten since launch

Symptom. The hint sets are the ones the account launched with. Nobody defends them; nobody has changed them either.

Cause. No owner and no cadence, which is the same cause as most maintenance failures anywhere. It is aggravated here by the absence of any prompt: there is no quality score turning amber, no disapproval notice, no “low search volume” flag. A stale hint set produces the same silent, unremarkable numbers it produced when it was fresh.

Fix. Put a recurring monthly rewrite window in a calendar with a name against it, and treat the weekly read as the input to it. Annotate every change with a date, because without annotations you cannot later separate your own edits from platform changes like the 17 August 2026 advanced-matching enablement, which moves reported conversion volume on its own.

6. One ad group carrying several intents

Symptom. Respectable CTR, poor conversion rate, and a persistent sense that no version of the ad copy works properly. Every rewrite improves one thing and worsens another.

Cause. The ad group violates one ad group, one intent. Somebody problem-aware and somebody comparing a shortlist need different headlines, different proof and different pages, and the ad group only has one of each. Since relevance is computed jointly from hints, landing page, title and copy, an ad group spanning two stages is guaranteed to be partly irrelevant to everybody it reaches, which shows up as a clearing price you did not need to pay.

Fix. Split it, and rewrite the ads so each new ad group answers one stage. Expect each arm to accumulate data at roughly half the old rate, so extend your read window before you judge the result — a split that looks like a regression at two weeks frequently is not.

7. Hints describing the buyer you want rather than the buyer you have

Symptom. The account converts, and the conversions are the wrong shape: smaller than the deals you priced the campaign around, or from segments your delivery team dreads, or from a market you cannot service well.

Cause. The hint set was written from the ideal-customer-profile slide rather than from closed-won data. It describes enterprise procurement language while every deal that actually closes is a twelve-person team, or the reverse. This one is worth catching early because the audience layer cannot rescue you from it: there is no age or gender targeting, no retargeting audience and no cross-device identity, and custom audiences require a minimum of 25,000 matched users with 100,000 or more recommended. Ad-group bid multipliers run 0.1x to 10x but never gate eligibility — 10x includes nobody new and 0.1x excludes nobody — so you cannot multiply your way to a different buyer.

Fix. Export closed-won by segment, and write the hint sets from the deals that actually closed and were actually profitable. Then price each ad group against its own maximum allowable CPA rather than an account average, since maximum allowable CPA is roughly average order value × gross margin × close rate and those three numbers differ by segment. If the buyer you want and the buyer you have are genuinely different people, they are two ad groups with two budgets, and one of them may deserve to be turned off.

What this means for your account

Two of these seven do most of the damage, and neither of them looks like a problem from inside Ads Manager. Hints that describe the product rather than the buyer’s moment produce an account that is fully configured, orderly and quietly underpriced against its own potential. Hints so broad that the ad groups overlap produce an account that cannot answer any question you ask it, for as long as it runs. The other five announce themselves eventually — through missing impressions, an obviously stale edit date, or ad copy that never quite works. Check those two first, in that order, and check them by reading the hint sets aloud rather than by reading the numbers.

Tools, LLMs, and what they can honestly do

A language model will generate a plausible set of context hints for almost any business in about a minute, and for a large number of advertisers that is the correct place to start. This is the strongest objection to paying anybody for hint research, it is a fair objection, and it deserves a straight answer rather than a defensive one. Generated hints are frequently good. They are structured, they use the buyer’s vocabulary more often than an internal marketing team does, and they cost effectively nothing to produce.

A category of hint-research and prompt-research tooling has also grown up around this channel, sold on ordinary monthly subscriptions at a small fraction of a strategy engagement, and we recommend that route without reservation where it fits.Below roughly $1,500/month in media, running ChatGPT Ads yourself is usually the right call, and at that spend a generator plus a careful hour is a better use of your money than any consultant, including us. We publish a free context hint generator for exactly that reason. Use it.

Three situations make generated hints plainly the right answer, and we say so on calls. You are launching a first test at $1,000–$3,000 and need a working account this week rather than an optimal one next quarter. You sell one product to one kind of buyer, so the architecture question that follows is close to trivial. Or you have a competent operator in-house who will do the weekly read and simply needs a faster way to draft candidates each month. In all three, buy the tooling and skip the engagement.

What does generation not solve?

Generation solves the blank page, which is a real problem worth solving. It does not solve four others, and the fourth is the one that decides whether this work has a person in it.

There is a fifth, quieter problem. A model generating a hint set for one ad group has no picture of what is already in your other ad groups, so it produces overlapping sets by default — failure pattern four, arrived at by the shortest available route. Generating five ad groups’ worth of hints in five separate sessions reliably produces five sets that cannot be told apart in the auction.

Where does the cost of being wrong actually sit?

Generating candidate language is cheap and getting cheaper every quarter, and anybody selling that as the deliverable is selling something whose price is falling toward zero. The expensive decisions are structural: which ad groups exist, what each one is allowed to contain, and what evidence would change your mind. Those three carry the cost of being wrong, because they are the ones you cannot cheaply reverse.

Consider the asymmetry. A weak hint costs you a slice of one ad group’s delivery for a few weeks, and you retire it at the next weekly read. A wrong architecture costs you the ability to read anything at all for the life of the account — overlapping ad groups produce inconclusive comparisons indefinitely, and you cannot tell that is what is happening from the inside. Some adjacent mistakes are worse still: a campaign objective is locked permanently at creation, and a mismatched conversion event does not backfill when you correct it, so the period it was wrong stays wrong forever. Those are the decisions worth paying somebody to get right, and they are all made before the first hint is written.

Which parts of the job can a tool do?

TaskA tool or LLM does this wellWhat still needs a personWhy
Producing candidate phrasings Yes — volume, variety and speed, far beyond what a person will patiently write. Selection. Most candidates are plausible and only some describe a conversation your buyers have. Plausibility and truth are different properties, and only one of them is observable from the outside of your business.
Translating product language into buyer language Yes, and usually better than an internal team too close to the product. Supplying the source material — calls, tickets, lost-deal notes — and judging which vocabulary is really yours. Without your own transcripts a model returns category-average language, which is exactly what every competitor also gets.
Drafting ad titles and copy per ad group Yes. Copy is the task generation is genuinely strongest at. Making title, copy, landing page and hints agree with each other in one direction. Those four are the auction’s relevance inputs. Copy drafted apart from the hints is drafted apart from the thing it has to match.
Deciding which ad groups exist No. A model has no view of the account as a whole and no reason to keep sets apart. The whole decision, made once, against the four buying stages and your own segments. Overlapping ad groups make every future comparison inconclusive, and nothing in the platform reports the overlap.
Deciding which intents you can afford No. The inputs are private financial data. Margin, close rate and average order value by segment, then a maximum allowable CPA per ad group. Two intents identical in language can differ by an order of magnitude in what a click is worth to you.
Reading the outcome No, and this will not change. The data does not exist to read. Inference from seven ad-group metrics, on clean seven-day blocks, with the trailing 48 hours discarded. The platform reports nothing per hint, so there is no per-hint result for any tool to learn from.
Deciding what would change your mind No. A model will write you a plausible success criterion for anything. Committing in advance to a threshold and a date, and honouring both. Criteria chosen after the numbers arrive are not criteria, and small samples will supply a flattering reading on demand.
Maintaining the set as the market moves Partly — regeneration is cheap, so producing fresh candidates monthly costs nothing. Knowing what actually changed: a competitor entering, your clearing price rising, or a platform default shifting under you. None of those three appear in a hint generator’s inputs, and two of them look identical in the reporting.

Read that table as a division of labour rather than a case against tooling. Four of the eight rows are things a tool does well, and one of them — producing candidate phrasings — is the task most people imagine is the entire job. It is the cheapest part of it. Start with the context hint generator, feed it your own call transcripts rather than your homepage, and you will have a serviceable set of candidates in an afternoon. What you will not have is a decision about which ad groups should exist, a per-segment ceiling on what a click is worth, or a rule written down in advance for what would make you stop. That is the part this engagement is, and if you can do it yourself with the generator and a careful afternoon, do that instead.

What you receive

The engagement produces a hint architecture you can operate, not a list of phrases you have to interpret. The difference matters: a spreadsheet of hints tells you what to paste into a text box, whereas an architecture tells you which ad groups exist, what each one is allowed to contain, and what evidence would justify changing it.

A loose drift of clear irregular glass fragments on the left of a white field resolving toward the right into four compact clusters of electric-blue glass, each fused into a single coherent form and throwing its own pool of light.
The output is not the hints. It is the small number of coherent clusters the hints resolve into, and the ad-group structure built around them.

The language corpus. Raw customer language harvested from your sales calls, support history, reviews and community threads, with the source recorded against each extract. You keep this. It is reusable for landing pages, email and sales enablement long after the hint sets have been rewritten.

The cluster map. That corpus resolved into a small number of distinct buyer situations, with the reasoning for each boundary written down — including the near-misses we merged and why. The merges are usually more instructive than the clusters.

The ad-group architecture. Which ad groups exist, which cluster each one serves, what intent stage it sits at, and what it is deliberately excluded from covering. This is the deliverable that makes the account readable, and it is the one software cannot produce because it depends on your margins and your budget rather than on language alone.

The hint sets. Five to fifteen per ad group, written from the corpus rather than from your marketing site, ready to paste in. With them, the rule for extending each set yourself — because you will need to, and you should not have to call us to do it.

The read plan. What signal each ad group is expected to produce, what would count as a genuine result versus noise at your budget, and the honest minimum run before anyone should act on a difference. Written before launch, deliberately, so the interpretation is not invented after the numbers arrive.

A working session. Ninety minutes with Tarun walking the architecture and arguing with it. Your team knows things about your buyers that no corpus contains, and this is where those get folded in.

What this engagement is not

It is not ongoing optimization — that is managed service, where hint sets get rewritten weekly. It is not an account build; if you have no account, setup and launch includes the first hint sets. And it is not a substitute for testing: an architecture is a hypothesis about your buyers, and the read plan exists precisely because the hypothesis might be wrong.

Price and scope

Context Hint Strategy & Research

From $3,999 one-off, fixed before we start

Language harvest and corpus, cluster map, ad-group architecture, hint sets for every ad group, a written read plan, and a ninety-minute working session. Typically two to three weeks, depending on how quickly source material can be gathered.

FactorEffect on price
One product or service line, one marketAt the floor
Several distinct product lines with different buyersAdds per line — each needs its own corpus and its own clusters
Multiple markets or languagesAdds — hint language does not translate, it has to be re-harvested
No existing source material at allAdds, or we decline — see below
An existing account with performance history to read againstReduces effort, and usually improves the output materially

One thing worth flagging about the last two rows. This engagement runs on raw material. A business with recorded sales calls, a support archive and a body of reviews gives us something real to work from. A business with none of those has only its own marketing copy, which is exactly the language the method exists to get away from. If that is your situation we will say so on the scoping call, and the honest recommendation is usually to spend a month recording calls first.

When this is the wrong purchase

Buy this when hint quality is genuinely your constraint, and when you have enough spend for ad-group-level reads to mean anything. That is a narrower set of situations than the offer might suggest, and the cases below are common enough to be worth naming.

Your budget cannot support the architecture

A hint architecture is only useful if each ad group receives enough traffic to be readable. Below roughly $1,500 a month you are choosing between a structure you can read and a structure that reflects your buyers, and the right answer is a single ad group, five hints, and a longer time horizon. That does not need a research engagement. Read the writing context hints guide and do it yourself.

Your measurement is broken

Hint research optimizes toward a conversion signal. If that signal is wrong, better hints will move your account confidently in the wrong direction, and you will have paid for the privilege. Fix the measurement first — either with a tracking implementation if you already know it is broken, or an audit if you are not sure.

You have not launched yet

Account setup and launch includes the first hint sets. Buying research separately before there is an account to put it in means paying twice for overlapping work, and we will point that out rather than take both fees.

You want hints written, not an architecture designed

Then this is disproportionate and we will say so. Use our free context hint generator, or a language model, or any of the tools that now do this commercially for a small monthly fee. Generating candidate language is genuinely cheap and getting cheaper. There is no honest way to charge four figures for it, and we are not going to pretend otherwise to make a sale.

You have no raw material

No recorded calls, no support archive, no reviews, no community presence. We would be writing hints from your website, which is the failure mode the whole method is designed to avoid. Start recording sales calls, come back in a month, and you will get a materially better output for the same money.

The honest summary of this offer

Most of what makes context hints work is published free on this site and reproducible by anyone who reads it. What is genuinely hard is deciding which ad groups should exist given your margins and your budget, and committing in advance to what evidence would change your mind. If those two decisions are the ones you want help with, this is the right engagement. If you want a list of hints, take the free tools — and take them with our blessing.

Sources and further reading

This page is the commercial companion to a much longer body of free material. If you only read one thing, read the guide rather than this page — it goes deeper on the theory and it is not trying to sell you anything.

A closing caveat that applies to every claim on this page. Because the platform reports nothing at hint level, no one — including us — can show you per-hint performance data, and any supplier who offers to is describing something the platform does not produce. Everything here is built on ad-group-level outcomes, on OpenAI’s own published statements about how hints function, and on the discipline of writing down what you expected before you look at what happened.

Context hint strategy FAQ

What are context hints in ChatGPT Ads?

Context hints are natural-language descriptions, set at ad group level, of the conversations, topics and situations where your product may be relevant. OpenAI states directly that they guide ad matching but are not exact-match keywords and do not guarantee delivery in any specific conversation. Our context hints guide is the fuller explainer, with fifty worked examples across ten verticals.

How are context hints different from keywords?

Keywords match strings; context hints match situations. Operationally the differences are larger than that suggests: there are no match types, no negative-keyword list, no keyword planner and no search-term report. You also cannot see which hint produced a result, so hints have to be managed as sets at ad-group level rather than as individually optimizable line items.

How many context hints should I use per ad group?

Five to fifteen. Launch with about five and watch impressions for five to seven days before extending. Adding hints beyond that range tends to make the ad group cover several intents at once, which destroys your ability to read the result — and reading the result is the only feedback the platform gives you.

Can I see which context hints are performing?

No, and this is the single most important thing to understand about the discipline. There is no search-term report and no per-hint performance data. You read outcomes at ad-group level and infer. Any supplier offering per-hint optimization is describing data the platform does not produce, and it is worth asking them which screen that number appears on.

Can I just use ChatGPT or a tool to write my context hints?

Yes, and for many businesses that is a reasonable place to start. Generating plausible hint language is genuinely cheap and getting cheaper, and our own context hint generator is free. What generation cannot do is decide which ad groups should exist given your margins and budget, harvest language from your recorded sales calls, or read an outcome the platform does not report. That is what this engagement is.

What do I actually get from a context hint research project?

A language corpus harvested from your own customer conversations with sources recorded, a cluster map resolving it into distinct buyer situations, the ad-group architecture built around those clusters, hint sets for every ad group, a written read plan stating what each ad group should produce and what would count as a result, and a ninety-minute working session with Tarun to argue with all of it.

How much does context hint research cost?

From $3,999 as a one-off, fixed in writing before the work starts, for one product or service line in one market. Several distinct product lines, multiple markets or multiple languages each add above the floor, because hint language does not translate — it has to be re-harvested from customers in that market.

How long does the engagement take?

Typically two to three weeks. The variable is almost always how quickly source material can be gathered on your side — call recordings, support exports, review data. The clustering and architecture work is fast once the corpus exists.

Can you do this without access to my ad account?

Yes for the research and architecture, which run on customer language rather than account data. Read access to an existing account materially improves the output, because prior ad-group performance is evidence about which clusters actually matter, but it is not a prerequisite.

What if I have no sales call recordings or reviews?

Then we would be writing hints from your marketing site, which is exactly the language the method exists to escape. We will say so on the scoping call and recommend you spend a month recording sales calls first. You will get a materially better output for the same fee, and we would rather wait than take money for a weaker result.

How often should context hints be updated?

Weekly iteration is the rhythm this channel rewards, and it is precisely what stops happening when the work sits on a busy internal to-do list. In practice: review weekly, rewrite when an ad group's read justifies it, and treat any hint set older than a quarter as suspect — language moves, competitors enter, and a set written in June describes a market that no longer exists by September.

Do context hints guarantee my ad appears in those conversations?

No. OpenAI is explicit that hints guide matching but do not guarantee delivery in specific conversations. Delivery is decided by the full conversational context and by ad relevance, of which your hints are one of four inputs alongside the landing page, the ad title and the ad copy. The image is explicitly not an input.

How do I stop wasted spend from bad context hints?

There is no negative-keyword list, so the two available instruments are writing a tighter hint and excluding a custom audience at campaign level. Note that custom audiences need a minimum of 25,000 matched users, which rules the second option out for many advertisers, and that bid multipliers never gate eligibility — a 0.1x multiplier excludes nobody. Tightening the hint is usually the only real lever.

Is this just keyword research with a new name?

No, and the tell is what happens to the output. Keyword research produces a list you bid on individually with match types and negatives. Hint research produces an architecture, because the platform will not report on individual hints and offers neither match types nor negatives. The unit you can actually optimize is the ad group, so the deliverable has to be the ad-group structure rather than the phrase list.

Why is my ad showing in irrelevant conversations?

Usually one of three things: a hint broad enough to describe several situations at once, an ad group carrying more than one intent so the system has conflicting signals about what it serves, or hints written to describe your product rather than your buyer's moment. All three are diagnosable from the ad group's own delivery pattern, and all three are fixed by tightening rather than by adding.

Should this be included in setup or management instead?

Often, yes, and we will tell you so. Account setup and launch includes the first hint sets, and managed service rewrites them weekly. This engagement makes sense as a standalone purchase when you already have an account running, your measurement is sound, and targeting specifically is the thing you believe is holding you back.

What is the minimum spend for this to be worth it?

About $1,500 a month. Below that you cannot support enough ad groups for an architecture to be readable, and the right answer is one ad group, five hints and a longer time horizon — which does not need a research engagement. Read the writing context hints guide and do it yourself.

Do you use my customer data, and how is it handled?

We work from material you provide — call recordings or transcripts, support exports, review data. We sign your NDA rather than insisting on ours, we use the corpus solely for your engagement, and none of the worked examples published anywhere on this site come from client material. Every example on this page and in the guides is illustrative and labelled as such.

Not sure targeting is your constraint?

Thirty minutes with Tarun. Bring your ad groups and he will tell you whether hints are genuinely the problem or whether it is measurement, structure or the offer — and if the free guide and generator would get you most of the way, he will say that instead of quoting you.

Book a scoping call
Tarun Kapoor, founder of Context Hints, seated at a wooden desk with a soft city light behind him.
Tarun Kapoor
Founder & CEO, Context Hints

Twelve years of media buying across GroupM, WPP, Ogilvy & Mather, and Neil Patel Digital. Has personally owned media for Nestlé, Sage, Qualcomm, Aetna, Weight Watchers, Chubb and Novotel.