Case studies >>
FactSet
How FactSet automated entity matching against official registry data
FactSet has built a model that automatically matches the companies in its database to official registry records on OpenCorporates, as part of a larger strategy to ensure their legal entity data coverage is updated and accurate.
It uses software to achieve what a researcher used to do by hand: search for a company, work through the candidates that come back, and identify the one official record that is genuinely the right match. Across several thousand test companies, it reaches match rates of up to around 70 percent, with almost no false positives.
— John Rieger, VP, Director of Entity Content Collection, FactSet
Verification and disambiguation at global scale
FactSet’s coverage is global, and the questions its researchers answer require a high level of precision.
- Is this a real legal entity? A name in a feed is not proof of registration. Confirmation has to come from an official source.
- Which company is this, exactly? Search “ABC Industries” and several businesses share the name. Researchers need to identify the right one using location, officers, and formation date.
- Is it still current? Companies change names, change status, and dissolve. A profile is only as good as the last time it was checked against the source.
- Is it a duplicate? With millions of records, distinguishing a genuinely new company from one already in the database is constant work.
Those questions have to be answered consistently, across every jurisdiction FactSet covers, at a volume no manual process can keep pace with.
The structural problem: home registrations and branches
One of the technical challenges can be demonstrated using FactSet itself.
FactSet is incorporated in Delaware. It is also registered to do business in Connecticut, Illinois, Florida and California, and its headquarters is in Connecticut.
Those two facts pull in different directions:
- The home registration in Delaware is the definitive legal record of where the company is incorporated, but Delaware registrations carry very little detail.
- The branch registration in Connecticut is far richer, holding the addresses and officer information that populate a company profile, and it corresponds to where the business actually operates.
For FactSet’s clients, headquarters is a foundational attribute used to screen, sort and aggregate, and it’s where most of a company’s officers sit. But the jurisdiction of incorporation still matters. The record that answers where is this company legally incorporated and the record that answers where is this company are frequently not the same record.
OpenCorporates links the two. Search results flag whether a record is a branch, and the detailed company endpoint points back to the home company’s jurisdiction code and company number.
Dual mapping. Rather than choosing between the home record and the branch, FactSet holds both. The branch matching the physical headquarters becomes the primary mapping for retrieving and refreshing firmographic data, while the home company identifier is retained alongside it as the answer to the incorporation question. Researchers select the branch matching the headquarters they already have in front of them, and the home company link comes back with it.
Coverage-led triage. OpenCorporates publishes fill rates for every jurisdiction, showing how often officers, registered addresses and industry codes are populated, alongside refresh rates showing how frequently each jurisdiction updates. FactSet uses both to target its collection: pursuing the enrichment that a given registry actually publishes, and aligning its own refresh cadence to the cadence of the underlying source.
The identifier layer. Every OpenCorporates record is keyed to the registry-assigned company number combined with a jurisdiction code. Records also carry additional identifiers including LEI, CIK and CAGE. Rieger describes these as “really nice tools for us to enhance our own mapping processes” when identifiers arrive from a third party.
Relationship data. The statements endpoint surfaces connections in both directions – parent companies and home companies upward, subsidiaries and branches downward – supporting FactSet’s work on corporate hierarchies and subsidiary collection.
Matching runs as a machine process. The model runs a name search, inspects each candidate to determine whether it is a branch, calls the company endpoint to retrieve the home company where one exists, then compares the assembled data against FactSet’s in-house records to select the best match. It applies a confidence threshold and, where candidates tie, uses signals such as active versus inactive status to resolve them. FactSet’s engineering team is developing its own variant on the same logic.
Accuracy where it counts. Match rates reach around 70 percent, and the accuracy of those matches is the metric FactSet optimizes for.
“There are few if any false positives – that’s what we’re really trying to avoid.”
— John Rieger, Vice President, Director of Entity Content Collection, FactSet
For a reference data business, that is the number that matters. A wrong match propagates into client-facing products, while an unmatched record simply waits for review.
Scale. FactSet has mapped over a million of its entities to OpenCorporates records, and is actively growing that universe. Usage rose sharply through 2026 as the model moved through development, and a second FactSet group has been onboarded to the API for its own work. Volume is expected to increase again as the model reaches production.
Records that stay current. With mappings established and automated, FactSet can return to a company on a cycle to confirm the profile still holds – that the name hasn’t changed and the entity is still active – rather than relying on a point-in-time check made when the record was created.
Overall data quality improvement. FactSet’s decision to utilise the OpenCorporates API was part of a larger strategy to ensure the highest possible quality of their legal entity data coverage, with a specific focus on freshness and accuracy.
Why primary-source data made this possible
Automated matching only works on a foundation you can interrogate, and on data authoritative enough to treat as the answer rather than as one opinion among several.
Because every OpenCorporates record comes from an official registry, carries provenance and a last-updated date, and is keyed to a stable identifier, FactSet could build logic that reasons about the data: branch or home, active or inactive, how fresh, how complete. Coverage and refresh rates are published per jurisdiction, so the model can be tuned to the realities of each registry rather than treating the world as uniform.
That is what makes near-zero false positives achievable. Confidence in an automated match depends entirely on the quality and transparency of what sits underneath it.
For any organization maintaining a large entity database, the pattern is the transferable part: map systematically against a primary source, target collection using published coverage, align refresh cadence to the source, and let the registry do the verification.
FactSet is a global financial data and analytics company serving the investment community. Its digital platform brings together proprietary financial data, client datasets, third-party sources, AI-powered capabilities, and flexible technology to deliver tailored solutions for the buy side, sell side, wealth management, private equity, and corporate markets.
Users rely on FactSet to research public and private companies, understand their businesses and markets, and make informed investment and business decisions. That makes the quality and accuracy of the underlying company data foundational. Before FactSet can publish anything about a company, it needs to be certain the company is a real, legally registered entity, know where it actually operates, and be confident it isn't holding two records for the same business.
That is the job of FactSet's entity content collection group: hundreds of researchers turning companies from client requests, investment activity and third-party feeds, into verified profiles.
Fresh, standardised, fully auditable information, underpinned by our Legal-Entity Data Principles, this is data you can trust.