In July 2023, Trellix’s Advanced Research Center published one of the better pieces of ransomware victimology anyone has done. They pulled 8,943 confirmed victims across 97 groups off criminal leak sites, enriched each one with sector, location, revenue and headcount, and published the distributions.
Look closely at the charts and you find something the write-up does not dwell on.
Company size was unknown for 27.87% of victims, revenue for 32.36%, sector for 27.28%. In all three, unknown beats every named category, including industrial at 24.61% of sectors and the $10 million to $50 million band at 21.93% of revenues.
Roughly one victim in three could not be resolved to an organization with attributes attached.
This is not a criticism of Trellix. Their method was better than most, they published their unknowns instead of quietly dropping them, and the analysis held up.
It is a criticism of the practice. Almost every ransomware report, breach roundup and risk dashboard has the same hole, and most do not show it to you.
Unresolved entities are not randomly distributed
If the unmatched third were a random sample you could ignore it. Drop the nulls, report on what is left, and the shape of the distribution survives.
Four reasons a name fails
It is not random. In our experience, records fail for four reasons, and each one correlates with something you are trying to measure.
| Why the record fails | The bias it creates |
|---|---|
| The organization is small | Small. A 40-person contractor exists in a state registry and almost nowhere else. Large enterprises resolve at near-100%. |
| It is not a company | Public sector. No company number, no incorporation record, no revenue filing. |
| The name is a trade name, an abbreviation or a typo | Non-English, and complex structures. A brand where the entity is a holding company, a division rather than a parent, dropped diacritics. |
| The entity has ceased to exist | The most severely affected. It did not survive, or it was acquired before the analysis ran. |
“City of Atlanta” is not a legal entity in any registry, and public bodies are among the most attacked organizations on earth.
Comparitech counted 187 ransomware attacks on government bodies in the first half of 2026, up from 165 in the previous six months. Roughly one a day.
Leak site posts, meanwhile, are typed by affiliates, often in a second or third language. The string was never meant to be a database key.
Which makes the bucket the opposite of a random sample
Put those together and the unmatched population is disproportionately small, public sector, non-English-speaking and structurally complicated.
Which is to say it is disproportionately made up of the organizations that ransomware crews target most and that defenders understand least.
Which means the conclusion is understated, not wrong
The finding
Trellix’s headline was that ransomware victims are mostly small and mid-sized businesses, not the household names in the headlines.
Most came from companies with 51-200 employees at 20.57%, then companies under 50 at 16.91%, with the percentages falling as size rises.
Now apply the bias
If the unresolvable third skews small, and it does, then the true share of small businesses is higher than reported.
Everyone in this field has been arguing that ransomware is a small-business problem while working from data that systematically undercounts small businesses.
Nobody publishes a match rate
The same exercise, three years later.
| Trellix, 2023 | Black Kite, 2026 | |
|---|---|---|
| Victims | 8,943 | 7,551 |
| Groups tracked | 97 | close to 300 |
| Classified by | sector, location, revenue, size | NAICS, HQ country, revenue |
| Revenue resolved | roughly two thirds | 6,000 of 7,551, roughly 80% |
Different years, different collection, different enrichment stack, and a twelve point spread in the thing that determines whether your chart means anything.
Nobody in this industry publishes match rate as a first-class metric. It should be table stakes: it is the single number that says how much to trust everything downstream.
The same failure, wearing three other costumes
Ransomware reports are the visible version of this problem because the output is a chart and the gap has a label on it. In the security products we work with, the gap is invisible.
| Where the miss happens | What arrives | What the miss looks like |
|---|---|---|
| Dark web and credential monitoring | A breach dump: emails, domains, sometimes a scraped name | The alert does not fire late. It does not fire at all. |
| Third-party and vendor risk | ”AWS”, “Amazon Web Services”, “Amazon Web Services EMEA SARL” and “AMZN”, arriving as four vendors | The vendor never enters the assessment queue, and the backlog improves, because the denominator shrank. |
| Insurance underwriting | A name and an address on a broker’s form. Classification drives pricing | Not an error. A mispriced policy, found at claim time, years later. |
In security, a miss is silent
The pattern across all four: in sales and marketing enrichment, a miss costs you a lead and you notice. In security, a miss is silent.
Nobody files a ticket saying “we failed to notify a client we could not identify.”
Missing data does not look like missing data. It looks like a smaller problem than you have.
And the unresolved records belong to the clients least likely to have a security team of their own.
How to report this honestly
Four practices, none costing more than a paragraph.
Publish your match rate, broken out by attribute. “Sector resolved for 94%, revenue for 71%” is more useful than any chart that follows it.
Report unknowns as a category rather than dropping them. Trellix did, which is why their 2023 numbers are still useful. Excluding nulls quietly turns an honest gap into a fabricated precision.
Run a sensitivity check. Reassign the unmatched population to your smallest matched decile and see whether the headline moves. If it does not, the finding just got stronger for free.
Publish the unresolved strings. They came off a leak site, so they are already public. Someone with a different stack will resolve a chunk and tell you. Nobody does it.
Where the answer actually lives
The data is not missing, it is scattered
For almost every name in that bucket, someone has published the answer. Rarely in one place, and often not in any index.
| The record that failed | Where the answer actually is |
|---|---|
| The 40-person contractor | A filing in one state’s registry, in that state’s format, behind a search form rather than a download |
| The city or the school district | A budget document, a council agenda, an audit report. PDFs on their own websites |
| The trade name or the abbreviation | A “formerly known as” line in a filing, or a local news story about the rebrand |
| The dissolved or acquired entity | A dissolution record, a deal announcement, a paragraph in the acquirer’s accounts |
Four publishers, four formats. A little is a structured feed. Most is a document, and much of it is indexed nowhere, because the source has no bulk export and answers one query at a time.
What it would take
Two ways to close that gap, and only two.
Hold all of it. Every registry in every jurisdiction, every filing, every rebrand, kept current. Nobody has that, and what was never machine readable cannot be held that way at all.
Or go and get it, per question: visit each source that might carry the answer, read what comes back, and follow the breadcrumbs.
A name change leads to an old filing. The old filing names a parent. The parent’s accounts name the company you started with.
That is research, not lookup, and for the hardest records it has been the only thing that works. The long tail has always been resolved by people, one browser tab at a time.
Which is the real reason the unknown bucket is so large. Most of this industry solved the part that indexes cleanly, and stopped at the edge of it.
The part that should bother you
Group-IB’s teardown of Qilin’s affiliate panel describes what the Targets section of the panel holds:
victim details, ransom amount, payment deadline, company revenue pulled from sources like Zoominfo, and the content of the ransom note.
Ransom-ISAC’s analysis of The Gentlemen, a different group, found it using ZoomInfo for victim revenue research. Negotiation transcripts reference a victim’s “$163 million revenue” and a “branch in Thailand”.
Group-IB also records that RansomHub affiliates “cross-reference stolen financial records with publicly available revenue data to determine what a victim can pay.”
The affiliate has no match rate to report
So the ransom demand is a function of a firmographic record. The affiliate pricing the extortion and the analyst building the victimology chart are querying the same category of database, from opposite ends.
The difference is that the affiliate is not constrained by a licence, a contract or a match-rate SLA, and does not have to explain to anyone why a third of their records came back unknown.
They resolve the entity. That is the whole job. It would be good if defenders were at least as rigorous about it.
If you are doing victim enrichment, vendor discovery or breach attribution at scale, I would like to compare notes on match rates and how you grade them. Or start with the entity resolution API.