All posts
entity-resolutioncybersecuritydata-quality

The biggest category in most ransomware reports is "unknown"

A third of the victims in the best public ransomware research could not be resolved. Those records are not missing at random.

SavvyIQ

· 8 min read

In July 2023, Trellix’s Advanced Research Center published one of the better pieces of ransomware victimology anyone has done (archived copy, because the argument below depends on it). They pulled 8,943 confirmed victims across 97 groups off criminal leak sites, enriched each one with sector, location, revenue and headcount, and published the distributions.

Look closely at the charts and you find something the write-up does not dwell on.

Company size was unknown for 27.87% of victims, revenue for 32.36%, sector for 27.28%. In all three, unknown beats every named category, including industrial at 24.61% of sectors and the $10 million to $50 million band at 21.93% of revenues.

Company size
unknown27.87%
51-200 employees20.57%
Revenue
unknown32.36%
$10 million to $50 million21.93%
Sector
unknown27.28%
industrial24.61%
Unknown against the largest named category, in each of the three charts. Trellix says it outright for revenue and sector. For company size they call 20.57% "most victims" while a larger bucket sits beside it unlabelled. Bars share one scale.

Roughly one victim in three could not be resolved to an organization with attributes attached.

This is not a criticism of Trellix. Their method was better than most, they published their unknowns instead of quietly dropping them, and the analysis held up.

It is a criticism of the practice. Almost every ransomware report, breach roundup and risk dashboard has the same hole, and most do not show it to you.

Unresolved entities are not randomly distributed

If the unmatched third were a random sample you could ignore it. Drop the nulls, report on what is left, and the shape of the distribution survives.

Why does a victim name fail to resolve?

It is not random. In our experience, records fail for four reasons, and each one correlates with something you are trying to measure.

Why the record failsThe bias it creates
The organization is smallSmall. A 40-person contractor exists in a state registry and almost nowhere else. Large enterprises resolve at near-100%.
It is not a companyPublic sector. No company number, no incorporation record, no revenue filing.
The name is a trade name, an abbreviation or a typoNon-English, and complex structures. A brand where the entity is a holding company, a division rather than a parent, dropped diacritics.
The entity has ceased to existThe most severely affected. It did not survive, or it was acquired before the analysis ran.

“City of Atlanta” is not a legal entity in any registry, and public bodies are among the most attacked organizations on earth.

Comparitech counted 187 ransomware attacks on government bodies in the first half of 2026, up from 165 in the previous six months. Roughly one a day.

Leak site posts, meanwhile, are typed by affiliates, often in a second or third language. The string was never meant to be a database key.

Which makes the bucket the opposite of a random sample

Put those together and the unmatched population is disproportionately small, public sector, non-English-speaking and structurally complicated.

Which is to say it is disproportionately made up of the organizations that ransomware crews target most and that defenders understand least.

Which means the conclusion is understated, not wrong

The finding

Trellix’s headline was that ransomware victims are mostly small and mid-sized businesses, not the household names in the headlines.

Most came from companies with 51-200 employees at 20.57%, then companies under 50 at 16.91%, with the percentages falling as size rises.

Now apply the bias

If the unresolvable third skews small, and it does, then the true share of small businesses is higher than reported.

Everyone in this field has been arguing that ransomware is a small-business problem while working from data that systematically undercounts small businesses.

Why does nobody publish a match rate?

The same exercise, three years later.

Trellix, 2023Black Kite, 2026
Victims8,9437,551
Groups97 represented in the victim setclose to 300 monitored
Classified bysector, location, revenue, sizeNAICS, HQ country, revenue
Revenue resolvedroughly two thirds6,000 of 7,551, roughly 80%

Different years, different collection, different enrichment stack, and a twelve point spread in the thing that determines whether your chart means anything.

Nobody in this industry publishes match rate as a first-class metric. It should be table stakes: it is the single number that says how much to trust everything downstream.

We are not exempt from that. Customers and prospects hand us a file of their hardest names and grade the results themselves, and across those files we come back at 90% or better, which is an aggregate and not the per-attribute breakdown this post is asking everyone for.

The same failure, wearing three other costumes

Ransomware reports are the visible version of this problem because the output is a chart and the gap has a label on it. In the security products we work with, the gap is invisible.

Where the miss happensWhat arrivesWhat the miss looks like
Dark web and credential monitoringA breach dump: emails, domains, sometimes a scraped nameThe alert does not fire late. It does not fire at all.
Third-party and vendor risk”AWS”, “Amazon Web Services”, “Amazon Web Services EMEA SARL” and “AMZN”, arriving as four vendorsThe vendor never enters the assessment queue, and the backlog improves, because the denominator shrank.
Insurance underwritingA name and an address on a broker’s form. Classification drives pricingNot an error. A mispriced policy, found at claim time, years later.

In security, a miss is silent

The pattern across all four: in sales and marketing enrichment, a miss costs you a lead and you notice. In security, a miss is silent.

Nobody files a ticket saying “we failed to notify a client we could not identify.”

Missing data does not look like missing data. It looks like a smaller problem than you have.

And the unresolved records belong to the clients least likely to have a security team of their own.

How to report this honestly

Four practices, none costing more than a paragraph.

Publish your match rate, broken out by attribute. “Sector resolved for 94%, revenue for 71%” is more useful than any chart that follows it.

Report unknowns as a category rather than dropping them. Trellix did, which is why their 2023 numbers are still useful. Excluding nulls quietly turns an honest gap into a fabricated precision.

Run a sensitivity check. Reassign the unmatched population to your smallest matched decile and see whether the headline moves. If it does not, the finding just got stronger for free.

Publish the unresolved strings. They came off a leak site, so they are already public. Someone with a different stack will resolve a chunk and tell you. Nobody does it.

Where the answer actually lives

If the data is not missing, where is it?

For almost every name in that bucket, someone has published the answer. Rarely in one place, and often not in any index.

The record that failedWhere the answer actually is
The 40-person contractorA filing in one state’s registry, in that state’s format, behind a search form rather than a download
The city or the school districtA budget document, a council agenda, an audit report. PDFs on their own websites
The trade name or the abbreviationA “formerly known as” line in a filing, or a local news story about the rebrand
The dissolved or acquired entityA dissolution record, a deal announcement, a paragraph in the acquirer’s accounts

Four publishers, four formats. A little is a structured feed. Most is a document, and much of it is indexed nowhere, because the source has no bulk export and answers one query at a time.

What it would take

Two ways to close that gap, and only two.

Hold all of it. Every registry in every jurisdiction, every filing, every rebrand, kept current. Nobody has that, and what was never machine readable cannot be held that way at all.

Or go and get it, per question: visit each source that might carry the answer, read what comes back, and follow the breadcrumbs.

A name change leads to an old filing. The old filing names a parent. The parent’s accounts name the company you started with.

That is research, not lookup, and for the hardest records it has been the only thing that works. The long tail has always been resolved by people, one browser tab at a time.

Which is the real reason the unknown bucket is so large. Most of this industry solved the part that indexes cleanly, and stopped at the edge of it.

The part that should bother you

Group-IB’s teardown of Qilin’s affiliate panel describes what the Targets section of the panel holds:

victim details, ransom amount, payment deadline, company revenue pulled from sources like Zoominfo, and the content of the ransom note.

Ransom-ISAC’s analysis of The Gentlemen, a different group, found it using ZoomInfo for victim revenue research. Negotiation transcripts reference a victim’s “$163 million revenue” and a “branch in Thailand”.

Group-IB also records that RansomHub affiliates “cross-reference stolen financial records with publicly available revenue data to determine what a victim can pay.”

The affiliate has no match rate to report

So the ransom demand is a function of a firmographic record. The affiliate pricing the extortion and the analyst building the victimology chart are querying the same category of database, from opposite ends.

The difference is that the affiliate is not constrained by a licence, a contract or a match-rate SLA, and does not have to explain to anyone why a third of their records came back unknown.

They resolve the entity. That is the whole job. It would be good if defenders were at least as rigorous about it.

If you are doing victim enrichment, vendor discovery or breach attribution at scale, we would like to compare notes on match rates and how you grade them. Or start with the entity resolution API.

Common questions

What percentage of ransomware victims cannot be identified?
Roughly one in three in the largest public study. Trellix could not resolve company size for 27.87% of its 8,943 victims, revenue for 32.36% and sector for 27.28%. In all three distributions, "unknown" was the largest single category.
Are unresolved ransomware victims a random sample?
No. Records fail to resolve for reasons that correlate with what the research is measuring, so the unmatched population skews small, public sector, non-English-speaking and structurally complicated. Dropping it does not leave the distribution intact.
Does the unknown bucket make the conclusions wrong?
It makes them understated rather than wrong. Trellix found victims are mostly small and mid-sized businesses. If the unresolved third skews small, and it does, the true share of small businesses is higher than reported.
What is a match rate in threat research?
The share of records a dataset could resolve to a real organization, ideally broken out by attribute. It is the single number that says how much to trust every chart downstream of it, and almost nobody in this industry publishes one.
Why are public sector victims so hard to resolve?
Business registries contain incorporated entities, and cities, counties, school districts and national agencies do not incorporate. "City of Atlanta" is not a legal entity in any registry, so it has no company number, no incorporation record and no revenue filing to match against.

Resolve your first entity in minutes

Sign up, get API keys instantly, and make your first call with $25 in free credit. One request returns a verified entity, drawn from 140+ official registries.

POST /v2/entity-resolution/async
{
  "request_id": "reqa_2ZUKmavJxCx4GHMnpJsc9",
  "status": "COMPLETED",
  "data": {
    "status": "matched",
    "confidence": 98,
    "type": "business",
    "entity": {
      "id": "siq_2ZUKocPbFCPLClZ5XtHlJ",
      "name": "Apple",
      "primary_legal_entity": {
        "name": "APPLE INC.",
        "jurisdiction": "California"
      },
      "website": "apple.com"
    }
  }
}