Regulatory operations teams manage thousands of documents across eCTD modules, and getting each one filed under the right heading is foundational to everything that follows — from building the submission to being ready for an inspection. Misfiled documents create rework, delay filings, and introduce compliance risk. Yet most organisations still sort documents by hand, or rely on rigid naming rules that fall apart when conventions vary across regions, CROs, or systems inherited from an acquisition.

A more effective approach combines the speed and predictability of fixed rules with the adaptability of software that learns. The three-layer approach described here reflects a pattern emerging across modern regulatory platforms — one that balances cost, accuracy, and operational control.

Why One Method Is Never Enough

Sorting by naming rules works well when documents follow predictable conventions. A file named m1-2-cover-letter.pdf maps cleanly to Module 1, Section 2. But regulatory teams routinely receive documents from external partners, CROs and legacy migrations where naming conventions are inconsistent or missing entirely. A clinical study report might arrive as CSR_FINAL_v3_reviewed.docx with nothing attached to say what it is.

Relying entirely on AI, on the other hand, introduces cost and waiting time that are hard to justify when six or seven documents in ten could be sorted instantly by their names. It also raises a fair question about consistency — a real concern in regulated environments, where an audit trail has to show that decisions were made the same way each time and can be explained.

The practical answer is a layered approach, where each stage handles the documents best suited to it.

Layer 1: Naming and Metadata Rules

The first layer applies fixed rules — typically forty or more naming patterns mapped to eCTD document types. When a document’s name, location or properties match a known pattern, it is filed instantly, with no AI involved and no cost. Confidence in this layer typically sits between 75% and 95%, depending on how specific the match is.

This layer handles the bulk of well-organised material: cover letters, module indexes, regional administrative forms, and anything that follows an established convention. For organisations with mature document practices, this layer alone may handle half to sixty per cent of incoming documents.

The key consideration is maintaining the rules. They should be version-controlled, testable, and extensible without a software release. Storing them as configuration rather than building them into the software lets regulatory operations teams add patterns themselves as new document types appear or regional requirements change.

Layer 2: Learning From Your Own Corrections

Documents that match no rule pass to the second layer, which learns from decisions your team has already made. When somebody files or corrects a document by hand, that decision is kept as an example. Future documents with similar characteristics — their text, their properties, their structure — are compared against that accumulated history.

This layer is particularly effective for patterns specific to your organisation. If a particular CRO consistently names its clinical study reports in a non-standard way, a single correction teaches the system to recognise that pattern from then on. The learning is incremental and specific to your organisation, so your results improve based on your own document history rather than somebody else’s.

The cost is moderate, and more importantly this layer creates a traceable feedback loop: every decision can be traced back to the human judgement that taught it.

Layer 3: AI Against the Full eCTD Taxonomy

Documents still unfiled after the first two layers go to an AI model. At this stage the system supplies the complete eCTD classification hierarchy alongside the document’s text and whatever properties are available. The model works out where the document belongs and returns a confidence level with its answer.

This layer handles the difficult cases: documents with ambiguous content, unusual types not seen before, or files that could plausibly sit in more than one place. It is the most expensive stage per document, and it only ever sees the ten to twenty per cent that genuinely need this kind of reasoning.

Cost controls matter here. A well-designed system routes these requests through a single point that enforces per-customer limits, tracks usage, and falls back to a cheaper model as budgets are approached. Reuse also plays a part — if a near-identical document was filed recently, the previous answer is reused rather than paying for the work twice.

Bulk Filing and How It Fits Your Workflow

The three-layer approach earns its keep during bulk work — legacy migrations, a CRO handing over a large batch, or integration after an acquisition, where thousands of documents arrive with nothing attached to say what they are. Documents are processed in the background, each layer applied in turn, and only the uncertain cases are put in front of a person.

A practical workflow looks like this:

  • Automatic pass: Rules and learned patterns file 80–90% of documents with no human involvement
  • Review queue: Low-confidence results are presented to regulatory professionals with a suggestion and a confidence level
  • Correction feedback: Those human decisions feed back into Layer 2, continuously improving accuracy

This approach respects the reality of regulatory operations: full automation is neither achievable nor desirable for compliance-critical work, but removing the manual effort on routine filing frees your team for the documents that genuinely need expert judgement.

Implications for Regulatory Operations Leaders

For Senior Directors evaluating AI classification, the layered approach addresses several common concerns:

  • Auditability: Every decision is traceable — to a rule, to a prior human correction, or to an AI judgement with the full question and answer recorded
  • Predictable cost: The expensive stage only ever sees what the cheaper stages could not resolve, keeping cost per document manageable at scale
  • Incremental adoption: You can start with rules alone and turn on later layers as confidence in AI-assisted filing grows
  • Freedom to change AI provider: The AI layer should not be tied to one supplier, so you can switch without rebuilding everything around it

Accuracy here directly affects submission timelines, audit outcomes, and how your team spends its week. A layered approach that combines fixed rules, learned patterns and AI judgement provides the reliability regulated environments require without sacrificing the adaptability that today’s document volumes demand.

About DnXT Solutions

DnXT Solutions provides cloud-native eCTD publishing, review, and regulatory compliance tools for life sciences companies. With 340+ submissions published and 20+ customers, DnXT is the regulatory platform purpose-built for speed and accuracy.