Skip to content

Normalize lone surrogates in documentation and symbols - #505

Open
mindaugasrukas wants to merge 3 commits into
sourcegraph:mainfrom
mindaugasrukas:fix/documentation-lone-surrogates
Open

mindaugasrukas wants to merge 3 commits into
sourcegraph:mainfrom
mindaugasrukas:fix/documentation-lone-surrogates

Conversation

@mindaugasrukas

@mindaugasrukas mindaugasrukas commented Sep 23, 2026 •

Copy link
Copy Markdown

Prevent protobuf serialization failures when TypeScript produces lone UTF-16 surrogates in documentation or generated symbol identifiers. This covers string literals such as '\ud800' and prototype property names:

export function C() {}
C.prototype = { '\ud800': 1 }

Use String.prototype.toWellFormed() to replace unmatched surrogates with �, preserving valid Unicode and existing textual escapes. Normalize identifiers centrally in ScipSymbol so definitions, references, and relationships use consistent values. Apply the same treatment to generated signatures, JSDoc, and module documentation.

Add es2024.string to TypeScript's lib configuration because the existing ES2022 declarations do not include toWellFormed(). Without it, TypeScript rejects the new calls as an unknown string method. This adds the required type declarations while keeping the compilation target at ES2022. The project's existing Node.js requirement already provides runtime support, so no polyfill or runtime upgrade is needed.

Regression coverage verifies protobuf serialization and matching property definitions and references. It also checks adjacent surrogates: \ud800\ud800\udc00\udc00 becomes �𐀀�, preserving the valid middle pair.

Validation

  • Build passed (yarn build).
  • Tests passed (yarn test).
  • Formatting checks passed (yarn prettier-check).
  • Lint passed (yarn eslint).

@mindaugasrukas

Copy link
Copy Markdown
Author

@eseliger Could you review this PR when you have a chance?

@christoph-sg christoph-sg left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for the PR! While reviewing I looked at a few more ways for lone surrogates to appear in the generated index. For example this code will produce a Symbol with a lone surrogate in it.

export function C() {}
C.prototype = { '\ud800': 1 };

Let me know if you'd like to extend this PR to cover generated symbols as well, or if you'd like me to take over. In general I'd probably go with String.toWellFormed() instead of trying to escape lone surrogates. They're very unlikely to appear intentionally in source code.

@mindaugasrukas mindaugasrukas changed the title Escape lone surrogates in symbol documentation Normalize lone surrogates in documentation and symbols Oct 6, 2026
@mindaugasrukas

Copy link
Copy Markdown
Author

Thank you for the PR! While reviewing I looked at a few more ways for lone surrogates to appear in the generated index. For example this code will produce a Symbol with a lone surrogate in it.

export function C() {}
C.prototype = { '\ud800': 1 };

Let me know if you'd like to extend this PR to cover generated symbols as well, or if you'd like me to take over. In general I'd probably go with String.toWellFormed() instead of trying to escape lone surrogates. They're very unlikely to appear intentionally in source code.

Thanks for the review! I’ve extended the fix to cover generated symbol identifiers and switched to toWellFormed() for both symbols and documentation. I also added regression coverage for your prototype example.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants