Here is a small, extremely common data quality problem with an unusually satisfying answer.

When a dossier is imported into a regulatory platform, manufacturer records get created from whatever the submission says. The submission gives you a name — “Patheon Manufacturing Services LLC” — in the section context of a Module 3 document. So you get a manufacturer record with a name and nothing else. No country. No city. No site address. No FDA Establishment Identifier. No DUNS number.

Multiply that by every site in every imported application and you have a master data set that looks populated and cannot answer a single useful question. Which of these records is the same physical site? Is the “Patheon” in this application the same establishment as the “Patheon Manufacturing Services” in that one? When a change impact assessment asks which filings name this site, all it has to work with is a name.

The information was in the submission all along

FDA already requires this. Form 356h — the application form in Module 1 of every NDA, ANDA and BLA — lists each facility with its FEI number, its DUNS number, its full address and its role in manufacturing the product.

So the authoritative, agency-facing statement of your site details sits in the same dossier as the manufacturer record that is missing them. It has been checked. It has been filed. Nobody was reading it.

Finding the forms

DnXT’s approach builds on your filed submissions being searchable. Because every document in every published sequence is indexed and knows what kind of document it is, finding “every Form 356h ever filed in this dossier” takes seconds rather than a manual trawl through folders.

Forms are read newest first, because the most recent filing carries the most recent list of facilities. Each one is read from the PDF’s own form fields, so the facility details come back as structured information rather than text scraped off a page.

Matching conservatively, on purpose

Matching a facility on a form to a manufacturer record is the step that can silently corrupt master data, so it is deliberately cautious.

Names are standardised before comparison and “doing business as” aliases are honoured, because the same establishment routinely appears under its legal name on the form and a trading name in the submission. What the matching will not do is match across two different sites of the same company. A manufacturer with plants in three countries will have three facilities whose names differ only by a city or a suffix — and getting that wrong writes one site’s FEI onto another site’s record, an error that then spreads into every impact assessment downstream, looking entirely plausible.

Facilities that match nothing are listed as gaps and never used to create new records. A Form 356h also lists testing laboratories, packagers and control sites that may legitimately have no manufacturer record in your platform. Creating records for them automatically would fill your master data with entries nobody curated.

Fill the blanks. Report the conflicts. Decide nothing.

The process fills empty fields and nothing else.

If a site already has an FEI entered by a data steward and the form carries a different one, that is not a conflict for software to settle. It is a genuine discrepancy between what somebody typed and what was filed with FDA — exactly the kind of thing a data quality review should surface. So it is reported and left alone. The rule across the platform is that something the software worked out never overrides something a person confirmed.

The whole exercise is a preview by default: it shows you what it would change and writes nothing until somebody with the right authority approves it. You run it, you read the proposal, you decide.

The case that proves the design

Run against real filed submissions, this flags something a name-based approach could never see: the same FEI number appearing under a different company name on a newer sequence.

Same facility identifier, different company, later filing. That is not a data error. That is a site that changed hands — a contract manufacturer’s plant sold to another organisation between one submission and the next, which is routine in the CDMO market and genuinely material to anyone assessing supply chain risk.

The manufacturer record keeps its name. The proposal states plainly that the site has been relisted under a newer owner, and a person decides what that means. An automatic rename would have quietly rewritten history. Silently skipping it would have lost the finding altogether.

Why read it back rather than type it in

Every regulatory organisation has a master data backlog, and every proposal to fix it involves somebody entering information. What gets entered then becomes a third version of the truth, sitting alongside what the vendor told you and what you filed with the agency.

Reading it back out of the filing inverts that. What you end up with is, by construction, what you told FDA — the version that matters when an inspector asks. And where the filing cannot tell you something, you get an explicit gap instead of a confident blank.

The general point is worth stating on its own: before adding a field for somebody to fill in, check whether the answer is already in something you filed. Very often it is.


DnXT builds eCTD publishing, submission planning, document management and dossier review software for regulatory operations teams. Book a demo to see site details read back from your own filed forms against your own submissions.