Can Waters/Agilent chromatograph PDFs export structured data directly?
The clear answer is no. Waters Empower and Agilent OpenLab/ChemStation PDF reports cannot export structured data directly. This article unpacks why — covering data flow, vendor business strategy, compliance, and the realistic alternatives.
The clear answer: no. Waters and Agilent chromatograph PDFs cannot export structured data directly.
This isn’t a technical limitation — it’s a deliberate industry design. We’ll examine it from four angles: the data flow, vendor business strategy, compliance considerations, and realistic alternatives.
1. How chromatography data flows
Raw data from a chromatograph follows this path:
Chromatograph → Workstation software (.raw / .D / proprietary) → Report (PDF / paper)
↓
Database (Oracle / SQL Server)
The key fact: structured data lives in the workstation database — not in the PDF. A PDF is a presentation format. Once data is rendered as text, lines, and images, the structure is gone.
2. Why don’t vendors offer “export from PDF” features?
Business strategy: API licenses are major revenue
Waters Empower and Agilent OpenLab both expose data interfaces (API/SDK) — but they’re paid features, and not cheap:
- Waters Empower LIMS Interface license: ~$30,000 – $80,000+ per year
- Agilent OpenLab Data Stream: comparable pricing
- Multi-vendor environments mean stacking multiple licenses
If PDF could export data directly, no one would buy the API license. The business logic is straightforward.
Data integrity considerations
From a GMP perspective, the PDF is the “tamper-evident record.” If the PDF itself could export data, how would the export be controlled? Selective exports? Vendors don’t want these compliance questions.
Format proliferation
Even within one vendor, report templates vary by version. Waters Empower 2 and Empower 3 templates differ — and that’s before customer-defined templates. Maintaining “export from PDF” across all variants would be expensive.
3. Alternatives compared
| Approach | API license needed? | Data integrity | Batch | Difficulty |
|---|---|---|---|---|
| Direct DB read | Admin-level access | Highest | Yes | High (needs IT + validation) |
| Workstation built-in export | No | High | No | Medium |
| ChromaParse | No | High | Yes | Low |
| Generic OCR | No | Low | Yes | Medium (heavy review) |
| Manual entry | No | Operator-dependent | No | Low (but very slow) |
4. ChromaParse: a no-license-required path
ChromaParse takes a simple approach: since the data has to go through PDF, let’s extract it from the PDF cleanly.
Supported report formats
- Waters Empower (assay, related substances, system suitability, …)
- Agilent ChemStation / OpenLab CDS
- Thermo Scientific Chromeleon
- Shimadzu LabSolutions
Key properties
- Batch processing — submit multiple PDFs at once
- Numerical fidelity — extracted values match the PDF byte-for-byte (no rounding)
- Source-trace — every value traces back to its location in the original PDF
- No license required — no need for vendor API licenses
Waters/Agilent PDFs not exporting structured data isn’t a bug — it’s a feature (for the vendor). For labs without an API license budget, ChromaParse is the most balanced trade-off across speed, accuracy, and compliance.
5. Compliance considerations
For pharmaceutical companies, any approach for extracting data from chromatography PDFs must consider GMP compliance:
- Data integrity — extracted data must match the original record exactly
- Audit trail — the extraction process must be traceable
- Validation — any tool used for data extraction must be appropriately validated
ChromaParse’s source-trace feature (click an Excel cell → jump to highlighted region in PDF) supports this, but specific validation has to follow your company’s SOPs.
References
- FDA 21 CFR Part 11 — Electronic Records, Electronic Signatures
- WHO Annex 5: Guidelines on data integrity
- Waters Empower Software Information — https://www.waters.com/
- Agilent OpenLab CDS Data Streamer — https://www.agilent.com/