Sector Mapping

Why the Sector Taxonomy Is the Most Defensible Part of a Market Intelligence Product

Priya Nair
Why the Sector Taxonomy Is the Most Defensible Part of a Market Intelligence Product

When we talk about what makes a market intelligence product difficult to replicate, the discussion tends to focus on data access, processing infrastructure, or team expertise. Those matter, but they are not the hardest part to reproduce. The hardest part is the taxonomy.

A taxonomy, in this context, is the set of decisions about how a market is partitioned into sub-segments. Which activities belong together? Where does one competitive group end and another begin? Which SIC codes should be treated as a single segment and which should be split further? These are judgment calls, and they have a compounding effect on everything the product does downstream.

Why taxonomy decisions are not recoverable from public codes

The UK Standard Industrial Classification system was last substantially revised in 2007. It was designed for macro-economic statistics, not for the operational distinctions that matter in PE deal work. SIC code 43120 covers "site preparation" and includes everything from groundwork contractors to demolition specialists to contaminated land remediation firms. These are three fundamentally different businesses serving different buyer types, with different cost structures and competitive dynamics. They appear in the same code because they share a surface-level activity description.

Anyone building a market intelligence product for the construction services sector faces a choice: accept the SIC categorisation as given, or build a secondary classification layer that subdivides the code into groups that actually compete with each other. The first option is fast and consistent; it is also nearly useless for identifying acquisition targets, because a deal team looking for groundwork contractors does not want to see demolition specialists alongside them in the output.

A secondary classification layer built from signal analysis, rather than code assignment, is what allows a product to make that distinction. The signals are filing disclosures (what equipment categories appear in the fixed asset base), employment patterns (what job titles dominate the workforce), and language analysis of company descriptions and contract awards. Each of these signals, taken alone, is imprecise. Combined across many companies over time, they allow you to draw defensible boundaries that the SIC system cannot draw.

The compounding effect of taxonomy quality

Why does taxonomy quality have such a large downstream effect? Because every subsequent operation in the pipeline inherits the errors in the partition.

Consider a ranking model. If the model is ranking "most likely acquisition target among UK groundwork contractors," it needs to know which companies are actually in that group. If the input set is contaminated with demolition firms and remediation specialists, the model is simultaneously ranking across three different competitive universes, and the output will reflect an average that applies accurately to none of them. A senior director reviewing the output will notice immediately that the list feels incoherent.

The same applies to signal extraction. Revenue bands derived from staffing patterns have different calibration across activity types. A demolition contractor with 40 employees typically operates at a very different revenue per employee ratio than a groundwork contractor of comparable headcount, because the capital and labour mix is different. If you apply a single calibration function to both, your revenue estimates will be systematically wrong for one or both groups. The error is not in the estimation method; it is in the failure to partition before estimating.

Why it takes time and why time is the main barrier

Building an accurate sub-segment taxonomy for a single sector requires iterating through real company data until the partition stabilises. This is not a theoretical exercise. You have to look at the actual companies that cluster together, check whether the clustering makes operational sense, manually investigate edge cases, and adjust the signal weights until the false membership rate drops to a level that does not corrupt the downstream analysis.

For a sector with strong internal differentiation, reaching a stable partition can take three to six months of active refinement. During that period, you are accumulating a dataset of anomalies: companies that signal membership in a sub-segment they do not actually belong to, companies that straddle two groups, and companies where the filing data is structurally misleading because the entity is a holding company for a group that operates across multiple sub-segments.

That anomaly dataset is where the real value accretes. It is not directly visible in the product output, but it is what separates a taxonomy that survives contact with a knowledgeable user from one that falls apart the moment someone who knows the sector asks why a particular company appears on a list.

Replication from the outside is difficult not because the information is hidden, but because the work is dense and domain-specific. A competitor could access the same public filing data. They could not access our refinement history or the accumulated decisions about edge cases in any particular sector. That history is what makes the taxonomy accurate at the level where it matters for deal teams.

The wrong way to evaluate taxonomy quality

A common approach to evaluating classification systems is to test recall and precision against a held-out set of manually labelled examples. This is useful, but it measures the wrong thing for this application. What a deal team needs is not high recall on a general sample; they need high precision on the companies that sit in the top quartile of any ranking. If the taxonomy misclassifies companies in the middle of the relevance distribution, the error is tolerable. If it misclassifies companies in the top tier, the output is actively misleading.

The evaluation approach we have found more useful is adversarial: present the output to someone with direct operational knowledge of the sector and ask whether any company on the list seems categorically wrong. Categorical errors, where a company clearly does not belong, are the failure mode that damages trust. Marginal errors, where a company is borderline, are typically discussable and do not undermine the output's credibility.

This is not to say that standard precision-recall metrics are useless. They are a useful baseline check. But they are not sufficient for the kind of quality that makes a product defensible in front of a partner who has spent fifteen years investing in a particular sector. That kind of quality is built through iteration with people who know the sector, not through statistical validation alone.

What this means for coverage expansion

Building an accurate taxonomy for a sector takes time, and that creates a genuine tension between coverage breadth and output quality. The honest position is that depth in a small number of sectors produces more useful output than thin coverage across many sectors. A well-partitioned taxonomy for UK industrial services is more valuable to a deal team focused on that space than a poor taxonomy that nominally covers industrial services plus healthcare services plus professional services.

We have chosen, deliberately, to prioritise depth. For the sectors we cover currently, the taxonomies have been through multiple refinement cycles and are validated against real sector knowledge. For sectors we add in future, we will go through the same process rather than launching with a first-pass partition that is likely to produce categorical errors in the output.

This is not a universal answer. For some applications, broad coverage with acknowledged imprecision is the correct trade-off. For the deal team use case, where the output is a short-list of twelve names that a human being will act on, imprecision in the taxonomy propagates directly into wasted time. The value proposition depends on the taxonomy being right, not on the taxonomy covering everything.

Put structured sector intelligence to work

Thema produces ranked target lists for UK private sectors from public record data. Request a sector map for your current coverage area.

Request access View pricing

More from the blog

Why Hiring Signals Predict M&A Readiness Better Than Revenue Filings Sector Mapping Is Not Target Screening: and the Difference Matters Building a Sector Knowledge Graph From Publicly Available Data