Skip to main content

6 posts tagged with "normalization"

View All Tags

PDF/A vs. Canonical PDF: What's the Difference?

· 6 min read
Napzoom Inc.

PDF/A and canonical PDF sound like they should solve the same problem. Both are about making PDFs more reliable. Both matter to compliance teams. Both show up when engineers start asking whether a document will still mean the same thing later.

But they are not the same thing.

PDF/A answers: Will this document remain renderable and self-contained for long-term preservation?

Canonical PDF answers: Will this input produce stable bytes, a stable hash, and a record of what changed?

PDFCanon vs. Adobe PDF Services: Specialist Normalization vs. General PDF APIs

· 4 min read
Napzoom Inc.

Adobe is the name most people associate with PDF. Adobe PDF Services is a broad API suite for creating, converting, extracting, combining, and manipulating documents.

PDFCanon is narrower by design. It does not try to be a general PDF toolkit. It solves one infrastructure problem: turn untrusted PDFs into byte-stable, auditable canonical documents.

PDFCanon vs. Glasswall: Deterministic PDF Normalization vs. CDR

· 4 min read
Napzoom Inc.

Glasswall and PDFCanon both handle PDFs in security-sensitive contexts. They answer different questions.

Glasswall is primarily a Content Disarm and Reconstruction platform. Its job is to neutralize risky file content before it reaches a user or system.

PDFCanon is a deterministic PDF normalization API. Its job is to turn untrusted PDFs into byte-stable canonical documents with stable SHA-256 hashes and machine-readable audit reports.

qpdf as a Service: When to Self-Host vs. Use an API

· 7 min read
Napzoom Inc.

qpdf is one of the best tools in the PDF ecosystem. It can repair structure, rewrite cross-reference tables, normalize object streams, decrypt files you are allowed to process, linearize documents, and expose a huge amount of the PDF's internal shape.

That makes it tempting to build your own PDF normalization service around it.

Sometimes that is exactly the right call. Other times, the hard part is not running qpdf. The hard part is proving that your output is deterministic, secure, auditable, and stable after the next toolchain upgrade.

Why Does My PDF Hash Change Every Time? A Field Guide for Engineers

· 9 min read
Napzoom Inc.

I spent a long time figuring out why sha256sum invoice.pdf returns a different hash every time my accounting software re-exports "the same" document.

Turns out this is a fundamental property of the PDF format - and it quietly breaks a number of real-world systems that depend on stable file hashes: deduplication, audit trails, content-addressed storage, integrity verification.