De-Identified MCP Answers: How Small-Group Rules Work
De-identified MCP answers withhold small groups, round counts and block subtraction. How k-anonymity, completed weeks and HIPAA de-identification apply to ad data.
A de-identified MCP answer is an aggregate built so nobody inside it can be singled out: groups below a minimum size are withheld, the remaining counts are rounded, and any figure that could be subtracted from another to expose a small group is withheld too. Under HIPAA, those rules still need Safe Harbor or an expert's documented determination behind them. Curve MCP applies them before a number reaches your AI assistant, counting people over completed weeks. A campaign showing 3 conversions last week can point straight at a patient, and Curve's BAA stops at the AI vendor you connect, so the answer itself has to be safe.
Why 3 conversions from one campaign can identify a patient
Picture a men's health clinic running a Google Ads campaign for testosterone therapy in one suburb. Last week's report shows 3 conversions from that campaign. It carries no names, emails or phone numbers.
The front desk knows who booked a first TRT consult last week, and anyone with CRM access can filter new leads by week. Say only three people booked that consult last week: the report just told anyone holding that list that all three came from this ad, and the campaign name says what care they wanted. With 40 bookings and 15 conversions, nobody can tell which 15.
HIPAA treats health information as individually identifiable when it identifies a person or when there is "a reasonable basis to believe the information can be used to identify the individual" (45 CFR 160.103). A tiny count tied to a service line and a single week can meet that test for anyone holding a second list.
MCP makes the second list easier to reach: an assistant connected to your ad data and your CRM in one session can do the join itself, just by being helpful. So the count has to be safe before it reaches the model, the same reason you cannot paste a patient funnel into ChatGPT.
Small-cell suppression and k-anonymity in plain terms
A cell is one number in a report: conversions for one campaign in one week, or people who reached one funnel step. Small-cell suppression means a cell below a minimum size is not shown as a number. It is replaced with a marker such as "suppressed" or left out.
The idea underneath is k-anonymity: every person behind a released figure should hide among at least k people who look the same in the data. For aggregate reports, that means no released count describes fewer than k people. HHS's de-identification guidance cites El Emam and Dankar's paper on k-anonymity.
What minimum group size public agencies use
HIPAA gives no number. HHS guidance says "there is no explicit numerical level of identification risk that is deemed to universally meet the 'very small' level." So look at agencies that publish small health counts for a living:
- CMS: no cell containing a value of 1 to 10 can be reported directly, so the smallest reportable count is 11. Zero is allowed. No value of 1 to 10 may be derivable from other reported cells, including through percentages.
- CDC WONDER: statistics representing one to nine deaths or births are suppressed, and rates based on fewer than 20 deaths are flagged as unreliable.
- Washington State Department of Health: suppresses counts below 10. It reports that NCHS moved to a minimum of 10 in 2011 after finding that suppressing only counts of 1 to 4 "failed to prevent disclosure."
Note the last one: a federal statistics agency tried hiding only 1 to 4 and abandoned it. A marketing report that hides only 1s and 2s sits below every benchmark here.
Count people, not visits
A threshold only means something if it counts the right thing. Twelve sessions can be two people, and fourteen conversion events can be one patient who submitted a form, booked and rebooked. The number that gets checked has to be distinct people across the whole reporting window, counted exactly. An estimate that is off by two can put a group of 9 on the wrong side of the line.
Rounding: useful blur, weak on its own
Rounding releases 240 instead of 243. It blurs exact gaps between large numbers, where small groups hide inside big ones. It does not replace suppression: a 7 rounded to 10 still describes a handful of people.
The failure that catches marketers is derived metrics. Cost per conversion is spend divided by conversions, so it can hand back the exact number you suppressed:
- Spend shown as $1,236 and cost per conversion shown as $412. Conversions are 3, whatever the conversion cell says.
- Visitors shown as 1,250 and conversion rate shown as 0.24%. Conversions are 3 again.
The fix is to compute every ratio from numbers that were themselves released, after rounding, and to drop the ratio whenever an input was withheld.
Complement and differencing attacks
Hiding one cell is not enough if the reader can rebuild it by subtraction. Marketing data offers three ways to do it.
The complement attack
A report shows campaign A with 40 conversions, campaign B with 25, campaign C as "suppressed", and a total of 68. C is 3. Secondary suppression (Washington's guidance describes it as avoiding "inadvertent disclosure through subtraction") hides the next-smallest part too, so the gap spans two cells. The simpler fix is never releasing a total over parts released one by one, and Curve MCP does both.
The subset difference
Some counts sit inside others. Every goal completer is also a visitor, and everyone at funnel step 3 also reached step 2. If step 2 shows 120 people and step 3 shows 117, three people dropped out, and that small group is exposed by two large numbers.
The rule: when one count sits inside another, the true gap must be zero or at least the threshold, otherwise the smaller count is withheld. The check runs on real counts before rounding, since the actual gap is what matters.
Differencing across queries
The third version uses two questions instead of two cells. Ask for "the last 30 days" on Monday and again on Tuesday, and the difference is one day in, one day out. Ask for conversions across all campaigns, then across all but one, and the gap is that one campaign. Each answer passes the threshold, but the difference between them does not have to.
So a connector's inputs matter as much as its outputs. Free-text filters and arbitrary date ranges let the user, or the model, design the subtraction. Even an exact count of hidden rows leaks: "1 campaign withheld" versus "0" tells a reader that a few people did something, which is the fact suppression was meant to hide.
Completed-week windows close the date gap
Rolling windows invite differencing because they move every day. Completed-week windows move once a week. A window that always ends on the last full Monday-to-Sunday week and spans whole weeks leaves nothing smaller than a week to subtract, and each week in the series must pass the threshold on its own.
The trade-off is freshness: the current week never appears, so figures can run up to a week behind a live dashboard. For weekly budget reviews that costs little, and the figure people most want, yesterday's bookings, is exactly the small group this design refuses to release. More on that design in why Curve MCP answers in weeks, not rows.
Safe Harbor vs Expert Determination for marketing aggregates
HIPAA offers two de-identification methods in 45 CFR 164.514(b), and they fit weekly ad reports very differently.
Safe Harbor: a checklist written for records
Safe Harbor removes 18 identifiers of the individual and of relatives, employers or household members. They include names, geography smaller than a state, every date element except the year, phone and fax numbers, email addresses, IP addresses, URLs, device IDs and "any other unique identifying number, characteristic, or code." The covered entity must also have no "actual knowledge" that what remains could identify someone.
Both parts strain against campaign reporting. Read strictly, a weekly count of people who booked rests on dates directly related to individuals, and Safe Harbor allows nothing finer than the year. And the testosterone clinic knows who its 3 bookers were, which is hard to square with the actual-knowledge condition.
Expert Determination: a documented risk judgment
Expert Determination asks a qualified expert to find that "the risk is very small that the information could be used, alone or in combination with other reasonably available information, by an anticipated recipient to identify an individual," and to document the methods and results. It suits aggregates because it can weigh thresholds, rounding, subtraction rules and the audience together.
"Anticipated recipient" does real work with MCP. A careful expert will treat the AI vendor processing the conversation as an anticipated recipient, and anything another connector in the same session can reach as "other reasonably available information." A determination evaluates a specific set of rules, so ask any vendor which version of its ruleset a determination covers and what happens when it changes. Our HIPAA-safe MCP server checklist has the other questions worth asking.
How Curve MCP applies small-group rules
Curve MCP works with any client that can connect to an MCP server, including Claude, ChatGPT and Cursor. It answers organization-level questions (visitors, sessions, goal completions and funnel steps) and reconciles your server-side campaigns: each campaign's platform-reported spend, clicks and impressions sit next to the conversions Curve's server-side tracking recorded, with cost per conversion by completed week. Each rule above maps to a control:
- People over the window. It counts distinct people, not visits, across the whole reporting window, and checks the threshold against those counts.
- Small groups withheld. A group below the minimum is withheld entirely and marked as below the threshold, never shown as a small exact number. The number of withheld rows is itself rounded the same way, and zero is withheld like any small count, so "nobody" and "a few people" look identical. The tool states its exact suppression and rounding rules to your assistant, so any answer can be checked.
- Counts and spend rounded. People counts, activity counts and spend are rounded, and cost per conversion is calculated from the rounded figures, so it cannot hand back a withheld count. Clicks and impressions come through as the ad platform reports them, because they are the platform's own totals, not counts of your patients.
- Subtraction routes closed. Where one number could be subtracted from another to reveal a small group, the smaller number is withheld. That covers funnel steps, goals against visitors and each week against its window.
- Completed weeks, fixed windows. Last week, or the last 4, 13 or 52 weeks. No custom date ranges, and never the current week.
- Fixed inputs, no visitor strings. The assistant picks from fixed choices, not free text, and no string a website visitor can set (UTM values, page paths, referrers) is ever returned.
- Approved names only. A goal, funnel or campaign name appears only if the organization approved it. Otherwise it reads "(label hidden)".
- A final guard that fails closed. Every response is checked for anything that looks like an email, phone number, ID, date or name. On a hit, or if the guard cannot run, the whole answer is blocked rather than cleaned.
The service is read-only, and its database role cannot read contact details, form answers or journeys. One link opens the full sent, accepted and matched view for a campaign inside Curve. A person-level question gets an "open in Curve" link too; it needs a Curve login, works once, expires quickly and carries no data, so the detail stays inside the platform your BAA covers.
The honest costs: no revenue or ROAS, no channel, device, region or page breakdowns, figures up to a week behind, and plenty of withheld rows for a small clinic. That is the ruleset working.
Frequently asked questions
Is a count of conversions PHI?
It can be. A count of 2 or 3 attached to a service-line campaign and a specific week can be tied to people by the clinic, its agency or anyone with the CRM. Group size and what the reader already knows decide it, not whether a name appears.
What minimum group size should a healthcare marketing report use?
HIPAA sets no number, and HHS says no single level universally qualifies. Public benchmarks cluster around 10 or 11: CMS will not report 1 to 10, and CDC WONDER and Washington's health department suppress below 10. Treat anything lower as a question your vendor must answer in writing.
Does rounding alone make a small count safe?
No. A small count rounded up is still a small group. Rounding blurs gaps between large numbers, and suppression is what protects small groups, so use both.
Why not just use Safe Harbor for ad reports?
Safe Harbor was written for records. Its date rule allows nothing finer than the year, and its actual-knowledge condition is hard to meet when your team can match a small count to patients. Expert Determination fits weekly aggregates better.
Why do some campaign numbers come back withheld?
Because the group behind them is below the minimum size, or because releasing them would let someone subtract their way to a small group. Cost per conversion is withheld along with its conversion count. For person-level detail, the assistant offers an "open in Curve" link instead.
Can an AI assistant re-identify people by combining connectors?
It can try if you give it the pieces, such as a CRM connector with contact access in the same session (GoHighLevel's v1 MCP server has a contacts_get-contacts tool, for example). Suppression, rounding and fixed windows are built so no released ad-side number describes fewer than a minimum number of people, and no two released numbers subtract to a smaller group. The rest is governance, and our guide to where MCP tool calls leak health data covers it.
Where to start
Start with the reports you already send. Pull last month's campaign report and look for any cell under 11 (the CMS line), any cost per conversion or conversion rate sitting on a small count, and any rolling window someone refreshes daily. Those leak small groups today, with or without AI. Our PHI-free client reporting template shows a safer layout.
Then decide what an assistant may see before anyone connects one; our guide to what to allow when connecting Claude to clinic ad data walks through it. Aggregate rules cannot fix raw data that already leaks upstream, so run the free compliance scanner to find tracking scripts, such as the Meta Pixel and Google Analytics, that may expose patient data. To see Curve MCP answer campaign questions with small groups withheld and counts rounded, book a demo. The rest of how Curve handles health data, BAA included, is at curvecompliance.com.
Reviewed September 2026. HIPAA citations are to 45 CFR 160.103 and 164.514. Suppression thresholds are as published by CMS, CDC WONDER and the Washington State Department of Health.
Stay Compliant. Scale Confidently.
Join healthcare innovators who trust Curve for HIPAA-compliant ad tracking.Launch in hours, not months. Your growth stack, now HIPAA-safe.
Book a free tracking audit