Accessibility & Inclusive Product Engineering

Accessible Document Remediation: Why PDFs Are Often Harder Than Web Accessibility

A company with a genuinely accessible website can still have years of inaccessible PDFs — annual reports, policy documents, forms, whitepapers — sitting behind that same site, invisible to a screen reader user in a way the website itself no longer is.

01

Why document accessibility gets overlooked

Web accessibility programmes typically focus on the website’s own templates and components,…

02

What makes PDF remediation genuinely harder

A PDF has no inherent reading order the way HTML does — the visual layout and the underlying…

03

A practical remediation approach

Start with an inventory: not every PDF needs the same priority.

Why document accessibility gets overlooked

Web accessibility programmes typically focus on the website’s own templates and components, which is the right starting point — but PDFs and other documents are usually produced by a completely different workflow (a different team, a different tool, often exported from Word or a design application) that the web accessibility programme never touches. The result is a website that passes WCAG audits sitting next to a document library that was never in scope.

No Inherent Reading OrderVisual layout & underlyingstructure can divergeComplex TablesMerged cells, multi-levelheaders are hard to tagEmbedded FormsField labels, tab order,sometimes JS validationScanned DocumentsNo text layer at allwithout OCR first
None of these have an HTML equivalent that’s this hard — which is why PDF remediation is routinely underestimated against a web accessibility budget.

What makes PDF remediation genuinely harder

  • A PDF has no inherent reading order the way HTML does — the visual layout and the underlying tag structure can diverge completely, and a screen reader follows the tag structure, not what the page visually looks like.
  • Tables, especially complex ones with merged cells or multi-level headers, are notoriously difficult to tag correctly for assistive technology, far more so than an equivalent HTML table.
  • Forms embedded in PDFs need explicit field labels, tab order, and sometimes JavaScript validation messaging — none of which carries over automatically from how the form looks.
  • Scanned documents (common for older archival material) have no underlying text layer at all until OCR is applied, and OCR output still needs manual review and correction for accuracy.

A practical remediation approach

Start with an inventory: not every PDF needs the same priority. High-traffic, high-consequence documents — forms people need to complete, policies that affect rights or benefits, anything linked prominently from the website — go first. Low-traffic archival documents can often be handled with a lighter-touch approach or, where feasible, replaced with accessible HTML versions rather than remediated PDFs at all, since HTML is usually easier to make and keep accessible than a PDF long-term.

The more durable fix, where it’s feasible: for documents that get frequently updated or are primarily consumed online, consider replacing the PDF with an accessible HTML page rather than remediating the PDF repeatedly. A PDF accessibility fix has to be reapplied every time the document changes; an accessible HTML template stays accessible across updates by design.

Building this into the content workflow, not just a one-time project

A one-time remediation project fixes the current backlog but doesn’t stop new inaccessible PDFs from being created tomorrow. The durable fix is training content authors on accessible authoring practices in the source tool (proper heading styles in Word or InDesign, alt text at creation time, accessible table structures) so accessibility is built in at export rather than retrofitted after the fact.

Frequently asked questions

Does WCAG apply to PDFs, or only to websites?

WCAG’s success criteria apply to PDF content where it’s used to convey information, and PDF/UA is a related, more PDF-specific accessibility standard. Both the EAA and most WCAG-referencing regulations generally expect documents linked from or central to a service to be accessible, not just the HTML pages themselves.

Can automated tools fully remediate a PDF?

Automated tools can tag basic structure and flag obvious issues, but reliable remediation — especially for complex tables, reading order, and forms — still needs human review. Automated tools are a good first pass, not a substitute for verification with actual assistive technology.

How do we prioritize a large backlog of PDFs realistically?

Rank by actual usage (page views, download counts) combined with consequence (does this document affect someone’s rights, benefits, or ability to transact) rather than attempting a full backlog in document-creation-date order, which rarely matches where the real user impact is.