Skip to content

How Should Pharma Handle Documents for AI?

5 min read

pharma document integrity AI

At the Industrial AI Summit 2026, attendees asked questions about pharmaceutical document integrity that the panel did not have time to answer live. Adam Procopio, Scientific Associate Vice President at Merck, and Kristen Sauter, President and General Manager of Life Sciences at Adlib Software, provided answers covering how to push data requirements into CDMO quality agreements, where to introduce traceability in the document lifecycle, and whether transforming documents into AI context introduces its own integrity risks. This is part one of two articles based on unanswered audience questions from the “You Closed the Sensor-to-Decision Gap. FDA Will Ask You to Explain It.” session.

Do You Push Document Standards Into the CDMO Agreement or Fix It After?

The original audience question: “Adam, you said you own the document accountability. But you also cannot control how a CDMO produces the record. So do you push your requirements into the quality agreement or accept the PDF and absorb the work? And where do you have technology doing it and where do you still have gaps?”

Adam Procopio, Scientific Associate Vice President at Merck:

We certainly want to push digital data requirements into our quality and technical agreements whenever possible because prevention is better than remediation. But the reality is we’re operating across a diverse CDMO ecosystem with varying levels of digital maturity. We cannot wait for every partner to produce perfectly structured data before we gain value from AI. Therefore, we need technology that can ingest spreadsheets, scanned records, PDFs, and other fragmented sources while maintaining traceability back to the original record. The biggest successes today are in extraction, search, and evidence linking. The biggest gaps remain semantic interpretation, cross-document reasoning, and generating conclusions that are consistently explainable and defensible in a regulatory environment.

Kristen Sauter, President and General Manager of Life Sciences at Adlib Software:

The technology piece Adam described (like ingesting spreadsheets, scanned records, fragmented PDFs while maintaining traceability) is exactly the problem the document accuracy layer is designed for. The key architectural requirement is that the transformation from whatever format you receive to a normalized, AI-ready output must itself be auditable. You need to be able to show: here is what we received, here is what the system did to it, here is what entered the AI workflow, and here is the confidence score on the extraction. The gap Adam identified (semantic interpretation and cross-document reasoning) is real and I’d add it’s partly a document quality problem, not just a model problem. Cross-document reasoning fails when the documents don’t share a common structure or vocabulary. Normalizing at ingestion, even imperfectly, reduces that failure surface significantly.

Where Should Traceability Start: At Generation or Receipt?

Kristen Sauter, President and General Manager of Life Sciences at Adlib Software:

I want to build on what I said in the session, because “build it at the source” is easy to say and hard to implement when your source is a CDMO with a different ERP, a different document management system, and a quality agreement written three years before anyone was thinking about AI. The principle holds: traceability retrofitted at receipt is always incomplete. When a batch record crosses a boundary (like CDMO to sponsor, manufacturing site to regulatory submission) you lose the generation context. You get a PDF or a scan, and the metadata that would tell you which version was approved, who signed it, what system produced it, is gone. You’re working from a copy of a copy. What actually works: push structured output requirements into quality and technical agreements, even when partners resist. Specify file format, required metadata fields, version identifiers, and handoff protocols. Not every CDMO will comply immediately, but having the requirement in writing changes the conversation, and protects you when something goes wrong. At the boundary itself, receipt processing should do two things the moment a document arrives: lock the version (so what you received is immutably recorded) and validate structure (so you know immediately whether the document is processable or needs remediation before it enters any downstream workflow). Doing that at receipt doesn’t replace traceability at source, but it does create a defensible handoff record. You can prove what you received, when, and in what state. The gap that most organizations still have: the connection between what was received and what the AI actually saw. If a document is transformed (normalized, chunked, embedded) and there’s no audit trail of that transformation, the traceability chain breaks exactly where regulators will look.

Does Transforming Documents for AI Break Things?

The original audience question: “What about transformation of docs into context for analytics? Does the transformation process break things?”

Kristen Sauter, President and General Manager of Life Sciences at Adlib Software:

Yes. And this is one of the most underestimated risks in enterprise AI programs right now. Every transformation step between the original document and the AI model is a potential integrity loss. Chunking splits tables across boundaries, so the column headers end up in one chunk and the values in another. Summarization compresses evidentiary detail (such as the exact specification range, the lot number, the reviewer’s qualification) into a general statement that sounds accurate but can’t be traced back to a specific claim. Embedding converts meaning into vectors, and the ranking that comes back depends on similarity, not on which document is authoritative. None of these transformations are wrong in principle. The problem is doing them without first making the document AI-ready in a way that preserves what matters: structure, hierarchy, field-level identity, version, provenance. What we mean by AI-ready is not “digitized.” A scanned batch record is digitized. An AI-ready batch record has its tables intact, its sections labeled, its critical fields extracted and validated, and its provenance locked before a single chunk is generated. When you transform a document that’s already structured and validated, the chunking strategy can respect document boundaries rather than imposing arbitrary ones. The embedding carries structured metadata. The retrieval is reproducible. The test I’d give any team evaluating whether their transformation process is safe: can you take an AI output, trace it to a specific chunk, trace that chunk to a specific passage in a specific version of a specific document, and show that the passage means what the AI said it means in context? If any step in that chain is opaque, the transformation broke something you’ll need later.

These questions were submitted by attendees of the “You Closed the Sensor-to-Decision Gap. FDA Will Ask You to Explain It.” session, at the Industrial AI Summit 2026. Answers provided by Adam Procopio of Merck and Kristen Sauter of Adlib Software. The session also featured Maya Schushan-Orgad of Teva Pharmaceuticals, and it was moderated by Rick Franzosa of Tech-Clarity. The session was sponsored by Adlib Software.

Related from IIoT World

Share with your network