Regulatory document libraries grow continuously. A mid-size pharmaceutical company may hold tens of thousands of documents across active and archived submissions, clinical study reports, correspondence and internal governance records. Finding the right document — or the right passage within a document — is a daily requirement that directly affects productivity, submission quality and inspection readiness.

Traditional keyword search has served this need for decades, but it fails in predictable ways: synonyms, terminology that varies by region, and plain-English questions all produce poor results. Meaning-based search addresses those limitations but brings trade-offs of its own. Combining the two — and merging the results intelligently — offers a more reliable answer for regulatory teams.

Where Keyword Search Falls Short

Keyword search excels at exact matches. Search for “adalimumab” and you get every document containing that word, ordered by how often it appears and where. For regulatory professionals who know exactly what they are looking for, it is fast, predictable and well understood.

The limitations appear as soon as the question gets more complex:

  • A search for “biosimilar approval pathway” will miss documents that say “abbreviated BLA” or “351(k) applications” — the same concept, different words
  • A plain-English question such as “what clinical endpoints were used in the Phase 3 trials for this application” returns noise, because keyword search matches individual words without understanding what is being asked
  • Terminology that differs by region (a US “drug master file” versus an EU “active substance master file”) splits results across what is really one document type

These limitations are not theoretical. Regulatory operations teams regularly report spending significant time hunting for documents they know exist but cannot find by keyword alone.

What Meaning-Based Search Adds

Meaning-based search converts documents and questions into a mathematical representation of what they are about, rather than which words they contain. Two passages discussing the same concept in different words end up with similar representations, so one can be found by searching for the other.

A question about “regulatory strategy for combination products” will surface documents discussing device-drug combinations, co-packaged products and cross-centre coordination, even where none of those exact words appear in the question.

For regulatory work, this is particularly valuable in a few situations:

  • Exploring a topic: when somebody is investigating an area rather than hunting for one known document
  • Finding connections: surfacing documents across different submissions that address related subjects
  • Asking questions in plain English: putting a question to the library rather than guessing which words it used
  • Mining legacy content: finding relevant material in archived submissions where naming and properties are inconsistent

It has its own weaknesses. It can return results that are related in subject but not actually useful. A search for “stability data for Product X” might surface stability protocols, stability reports for a different product, and general guidance on stability testing — all on topic, not all wanted. It also requires more setup and ongoing upkeep than keyword search alone.

Combining the Two

The better answer runs both searches at once and merges the two ranked lists.

The merging method that works best here scores each result by where it placed in each list, adds those scores together, and re-orders. A document near the top of both lists ends up highest overall. A document that ranked well in one but is absent from the other still appears, further down. Crucially, this works without having to compare two scoring systems that were never meant to be compared with each other.

The merge preserves the strengths of both:

  • Exact matches — drug names, document numbers, regulatory form numbers — are caught by the keyword side
  • Concepts and plain-English questions are caught by the meaning-based side
  • Documents that satisfy both — containing the right words and being genuinely on topic — rise to the top

Choosing the Right Method Automatically

Not every search needs both. Looking up document number “m1-2-3-cover-letter-12345” is a pure keyword lookup, and adding meaning-based processing would only slow it down. A question such as “how did the agency respond to our CMC deficiency” is the opposite: matching individual words would only add noise.

A good search system works out which kind of question it has been asked:

  • Exact lookup: document numbers, phrases in quotation marks, form numbers, or specific property values
  • Meaning-based: plain-English questions, exploratory searches, conceptual enquiries
  • Both: questions that mix specific terms with general concepts
  • Direct answer: questions that should return an answer drawn from the documents rather than a list of documents

Left on automatic, the system chooses for each search, which takes the decision away from the user entirely. Most regulatory professionals should never have to think about search modes — they type their question and get relevant results.

Narrowing Results the Way Regulatory People Think

Results benefit from filters that match how this profession actually organises work:

  • eCTD Module: narrow to Module 1 through Module 5 for administrative, clinical or quality content
  • Region: isolate results for the US, EU, Japan or Canada
  • Application: restrict to a specific product application, NDA or BLA
  • Document type: clinical study reports, correspondence, labelling, specifications, and so on
  • Compliance domain: for organisations managing several disciplines, filter to regulatory, quality, clinical or another governance area

These filters work with both kinds of search, narrowing what comes back without disturbing how results were ranked.

What It Takes to Run

Running both kinds of search means maintaining both: a conventional word index and a meaning-based one. Each customer’s content is kept entirely separate, so results can never cross organisational boundaries — a hard requirement for any shared regulatory platform.

The meaning-based side does its work when a document arrives: the text is extracted, broken into passages, converted into its mathematical form and stored. Regulatory documents are mostly text-heavy PDFs, which suits this well. Libraries in the tens of thousands of documents remain comfortably within what modern systems handle.

Responsiveness benefits from reusing recent work. Repeating a search does not repeat the underlying computation, and results for common searches are held briefly — typically between five and sixty minutes. For teams searching repeatedly while assembling a submission or preparing for an inspection, that makes a noticeable difference.

Practical Implications

For regulatory operations leaders evaluating search, this addresses a persistent and expensive frustration: time spent finding documents that exist somewhere in the library but cannot be found by keyword. Combining exact matching, understanding of meaning, and filters built around eCTD gives a search experience that matches how regulatory professionals actually think — sometimes by exact identifier, sometimes by concept, and most often by a mixture of both.

The investment is mostly in the meaning-based half. Organisations already running keyword search can add it gradually, starting with the collections where it pays off soonest — active submissions and recent correspondence — and extending to the archive once the benefit is obvious.

About DnXT Solutions

DnXT Solutions provides cloud-native eCTD publishing, review, and regulatory compliance tools for life sciences companies. With 340+ submissions published and 20+ customers, DnXT is the regulatory platform purpose-built for speed and accuracy.