Skip to content

BrsFile ast toString() - #754

Open
TwitchBronBron wants to merge 4 commits into
v1from
ast-to-string
Open

TwitchBronBron wants to merge 4 commits into
v1from
ast-to-string

Conversation

@TwitchBronBron

@TwitchBronBron TwitchBronBron commented Dec 2, 2022 •

Copy link
Copy Markdown
Member

Adds the ability to call .toString() on any BrightScript/BrighterScript AST node and get the source code it represents, including all leading trivia (whitespace, comments, newlines, colons). For an unmodified AST, ast.toString() returns exactly the text that was parsed.

This PR has been rebuilt on top of v1, which already tracks trivia on every token, so most of the original v0 work (lexer trivia, leadingWhitespace) is no longer needed.

const parser = Parser.parse(text);
parser.ast.toString() === text; // true

// works on any node, and reflects changes made to the AST
assignment.tokens.name.text = 'firstName';
parser.ast.toString(); // `firstName = "bob"`

How it works

  • AstNode has a new abstract toSourceNode(state: TranspileState): SourceNode, implemented by every statement and expression. toString() calls it with a blank TranspileState. Because it returns a SourceNode, the output carries source-map positions.
  • New TranspileState helpers:
    • tokenToSourceNodeWithTrivia
    • nodeToSourceNode (also writes any annotations)
    • nodesToSourceNode (writes a list of nodes with their separators)
    • statementsToSourceNode

Parser changes

Some tokens were consumed by the parser but never stored, so they are now kept on the AST:

  • tokens.commas on CallExpression, FunctionExpression, ArrayLiteralExpression, IndexedGetExpression, IndexedSetStatement, DimStatement, CallfuncExpression, InlineInterfaceExpression, TypedFunctionTypeExpression, and InterfaceMethodStatement. commas[i] is the comma after item i.
  • tokens.expressionBegins and tokens.expressionEnds (the ${ and } tokens) on template strings.
  • tokens.dot for the obj.[index] syntax on IndexedGetExpression and IndexedSetStatement.
  • tokens.eof on the root Body. It holds everything after the last statement.
  • Colons that are part of the AST (AA key: value, ternary a ? b : c, label:) are removed from the next token's leading trivia, so they aren't written twice.
  • When a file has syntax errors, any token the parser skipped is moved into the leading trivia of the next token in the AST. This means even broken code round-trips, which matters for the language server. The extra pass only runs when the parse produced diagnostics, so valid files pay nothing for it.

Plugin-created nodes

Nodes and tokens created by plugins usually have no location and no trivia. Written as-is, they would run into their neighbors (x=5+3, statements on the same line). So when a token or node has no location and no leading trivia, toString() adds sensible defaults:

  • spaces around operators, keywords, = and as
  • , between arguments
  • statements start on a new line

Tokens and nodes that came from the parser are never changed.

Bugs fixed along the way

  • TypedFunctionTypeExpression had its left and right parens swapped.
  • Cloning an AAMemberExpression dropped its comma, and cloning an AnnotationExpression dropped its call arguments.
  • exitwhile lost its leading trivia when it was split into exit + while.
  • The parser looped forever when #error was the last line of a file with no trailing newline.
  • The lexer included an unexpected character in the text of the next token, e.g. a newline token with text "%\n".
  • Characters the lexer doesn't recognize (e.g. a stray |) are now kept as TokenKind.UnexpectedCharacter leading trivia instead of being dropped, so they round-trip. They are never added to the token list, so parsing, diagnostics, and transpile output are unchanged.
  • When end sub or end class was missing at the end of a file, the previous token was stored a second time as the end token.
  • createToken(TokenKind.ForEach), createToken(TokenKind.ExitWhile), and createToken(TokenKind.At) now default to for each, exit while, and @.
  • A function with a sub/function token but no end token was transpiled (and typedef'd) as end [object Object].
  • TranspileState.sourceNode() crashed when given an undefined locatable.
  • ExitStatement.transpile crashed without a loop type token, and plugin-created exit/continue loop types were not separated by a space.
  • InterfaceMethodStatement.leadingTrivia crashed when the function type token was missing.
  • Typedefs dropped comments above an annotation (they're part of the annotation's @ token trivia), and moved comments between an annotation and its statement above the annotation. A new BrsTranspileState.getTypedefLeadingCommentsAndAnnotations writes both in source order for every typedef.
  • MethodStatement.leadingTrivia skipped the public/override modifiers, so comments above those methods were dropped from typedefs and from the transpiled method function.

Verification

  • The full test suite passes, and new specs cover every node type, clones, plugin-created nodes, the new parser tokens, and syntax-error recovery.
  • Coverage: every line and branch added by this PR is covered (100% of statements and branches in the changed source files). AstOutput.spec.ts also covers previously untested paths in the existing transpile and getTypedef methods, including SOURCE_NAMESPACE_NAME and SOURCE_NAMESPACE_ROOT_NAME, typedef leading comments and annotations, and nodes missing optional parts. The only uncovered branches left in those methods are defensive fallbacks that can't be reached (for example getLeadingComments(...) ?? [], since it always returns an array).
  • Every parse performed by the full test suite was checked with ast.toString() === input, and every unique input matches exactly.
  • 2604 real-world .brs/.bs files (rooibos, the Roku SDK samples, and a large production app) all round-trip exactly.
  • Transpiled output for 1365 real-world files was compared before and after this change. Ignoring comments and blank lines, the code is identical in every file. The only difference is 496 comment lines that used to be dropped and are now kept, mostly doc comments above public/override methods.

Known limitations

  • Exact fidelity requires locations, which is the default (trackLocations: true). With trackLocations: false, every token looks like it was created by a plugin, so default spacing may be added. trackLocations: false is only used in tests.

🤖 Generated with Claude Code

Every AST node now has a `toSourceNode(state)` method and a `toString()` method that
rebuild the node's source code, including all leading trivia (whitespace, comments,
newlines, colons). For an unmodified AST, `ast.toString()` produces the exact text that
was parsed.

- Add `tokenToSourceNodeWithTrivia`, `nodeToSourceNode`, `nodesToSourceNode`, and
  `statementsToSourceNode` helpers to `TranspileState`
- Store tokens the parser previously discarded: commas (call args, function params,
  array literals, indexes, dim, callfunc, inline interfaces, typed function types),
  template string `${`/`}` tokens, the `.` in `obj.[index]`, and the eof token on the root body
- Colons that are part of the AST (AA members, ternary, labels) are removed from the
  leading trivia of the next token so they aren't written twice
- When there are syntax errors, tokens skipped by the parser are moved into the leading
  trivia of the next token so the AST still represents the full source
- Plugin-created nodes/tokens (no location, no trivia) get sensible default spacing
- Fix swapped left/right parens in TypedFunctionTypeExpression
- Fix AAMemberExpression and AnnotationExpression clones dropping the comma and call args
- Fix `exitwhile` losing its leading trivia
- Fix infinite loop in `consumeUntil` when `#error` is the last line of a file
- Fix lexer including unexpected characters in the text of the next token
- Fix duplicated end tokens when `end sub`/`end class` is missing at the end of the file

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TwitchBronBron
TwitchBronBron changed the base branch from master to v1 September 26, 2026 19:45
@TwitchBronBron
TwitchBronBron marked this pull request as ready for review September 26, 2026 19:45
TwitchBronBron and others added 3 commits September 26, 2026 15:17
Add tests for the less common paths of every AST output method (`toSourceNode`/`toString`,
`transpile`, and `getTypedef`): nodes that are missing optional parts (as created by plugins
or by the parser during error recovery), and output methods that are normally bypassed by
their parent. Every line and branch added for `toString()` is now covered.

Fixes found along the way:
- `end [object Object]` was written for functions that have a `sub`/`function` token but
  no end token (transpile and typedef)
- `TranspileState.sourceNode()` crashed when given an undefined locatable
- `ExitStatement` crashed when transpiled without a loop type token
- Plugin-created `exit`/`continue` loop type tokens were not separated by a space
- `InterfaceMethodStatement.leadingTrivia` crashed when the function type token was missing
- `createToken(TokenKind.At)` produced `at` instead of `@`
- Plugin-created functions with no `sub`/`function` token got a leading space before the name

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Typedefs dropped comments that were above an annotation (they belong to the annotation's
  `@` token, not the statement), and wrote comments that were between an annotation and its
  statement above the annotation. Add `BrsTranspileState.getTypedefLeadingCommentsAndAnnotations`
  to write both in source order, and use it for every typedef that writes comments
- `MethodStatement.leadingTrivia` skipped the modifiers (i.e. `public`, `override`), so comments
  above them were dropped from typedefs and from the transpiled method function

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The lexer reported characters it doesn't recognize (i.e. a stray `|` or `%`) and then dropped
them, so `ast.toString()` couldn't reproduce them. They are now kept as
`TokenKind.UnexpectedCharacter` tokens in the leading trivia of the next token. They are never
added to the token list, so the parser, diagnostics, and transpile output are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant