Skip to main content

3 posts tagged with "pdf"

View All Tags

PDF/A vs. Canonical PDF: What's the Difference?

· 6 min read
Napzoom Inc.

PDF/A and canonical PDF sound like they should solve the same problem. Both are about making PDFs more reliable. Both matter to compliance teams. Both show up when engineers start asking whether a document will still mean the same thing later.

But they are not the same thing.

PDF/A answers: Will this document remain renderable and self-contained for long-term preservation?

Canonical PDF answers: Will this input produce stable bytes, a stable hash, and a record of what changed?

qpdf as a Service: When to Self-Host vs. Use an API

· 7 min read
Napzoom Inc.

qpdf is one of the best tools in the PDF ecosystem. It can repair structure, rewrite cross-reference tables, normalize object streams, decrypt files you are allowed to process, linearize documents, and expose a huge amount of the PDF's internal shape.

That makes it tempting to build your own PDF normalization service around it.

Sometimes that is exactly the right call. Other times, the hard part is not running qpdf. The hard part is proving that your output is deterministic, secure, auditable, and stable after the next toolchain upgrade.

Why Does My PDF Hash Change Every Time? A Field Guide for Engineers

· 9 min read
Napzoom Inc.

I spent a long time figuring out why sha256sum invoice.pdf returns a different hash every time my accounting software re-exports "the same" document.

Turns out this is a fundamental property of the PDF format - and it quietly breaks a number of real-world systems that depend on stable file hashes: deduplication, audit trails, content-addressed storage, integrity verification.