Skip to content

pdf_io: write the formatted buffer in pdf_dev_begin_actualtext - #1401

Open
peisbacher wants to merge 1 commit into
tectonic-typesetting:masterfrom
peisbacher:fix-actualtext-buffer
Open

peisbacher wants to merge 1 commit into
tectonic-typesetting:masterfrom
peisbacher:fix-actualtext-buffer

Conversation

@peisbacher

Copy link
Copy Markdown

pdf_dev_begin_actualtext formats each character into the local buf, but hands
work_buffer to dev_out. That buffer is never written in this function, so the
/ActualText string in the output consists of whatever work_buffer happens to
hold, in practice a run of NUL bytes.

Upstream texk/dvipdfm-x/pdfdev.c in TeX Live passes buf here. The divergence
came in with 3f16760 ("crates/pdf_io: update for TeXLive 2021 xdvipdfmx"), which
rewrote several pdf_doc_add_page_content(work_buffer, len) calls into
dev_out(p, buf, len) and changed only the sprintf target in this one.

Reproducing

\documentclass{article}
\usepackage{fontspec}
\setmainfont{Fira Sans OT}
\XeTeXgenerateactualtext=1
\begin{document}
Philipp
\end{document}

The seven characters produce seven NUL bytes:

tectonic 0.17.0:   /Span << /ActualText (\0\0\0\0\0\0\0) >> BDC
xelatex (TL 2022): /Span << /ActualText (Philipp) >> BDC

Why it is worth fixing

/ActualText takes precedence over the font's ToUnicode CMap in poppler, so a
document built with \XeTeXgenerateactualtext=1 has no usable text layer at all:
pdftotext returns NUL bytes for every glyph in the affected spans. Copy-paste,
search, screen readers and CV/resume parsers all see nothing. Extractors that
ignore /ActualText, such as pypdf, are unaffected, which makes this easy to miss.

The flag exists precisely to correct a wrong ToUnicode mapping, so the documents
that set it are the ones that can least afford a broken text layer. In my case the
font maps U+002D, U+00AD, U+2010 and U+2011 to one glyph and xdvipdfmx picks the
highest codepoint for the reverse mapping, which put U+2011 where hyphens belonged.

Verification

Built this branch on macOS 27 (aarch64) and ran the snippet above:

release 0.17.0:  /Span << /ActualText (\0\0\0\0\0\0\0) >> BDC   pdftotext: \0\0\0\0\0\0\0
this branch:     /Span << /ActualText (Philipp) >> BDC          pdftotext: Philipp

All three encoding paths of the function check out:

input /ActualText bytes pdftotext
Philipp Philipp Philipp
Größe Öl Straße Gr\xf6\xdfe, \xd6l, Stra\xdfe (PDFDocEncoding) Größe Öl Straße
C‑Level (U+2011) \xfe\xff\000C \021\000L... (UTF-16BE) C‑Level

Happy to add a regression test if you would like one, though I did not see an
existing pattern for asserting on content-stream output.

The function formats each character of the /ActualText string into the local
`buf`, but hands `work_buffer` to dev_out(). Nothing in this function writes
work_buffer, so the string that reaches the PDF is whatever that shared scratch
buffer happens to hold, in practice a run of NUL bytes.

Upstream texk/dvipdfm-x/pdfdev.c passes `buf` here. The divergence came in with
3f16760, which rewrote several pdf_doc_add_page_content(work_buffer, len) calls
into dev_out(p, buf, len) and changed only the sprintf target in this one.

Because /ActualText takes precedence over the font's ToUnicode CMap in poppler,
a document built with \XeTeXgenerateactualtext=1 ends up with no usable text
layer at all: pdftotext returns NUL bytes for every glyph in the affected spans.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants