BrsFile ast toString() - #754
Open
TwitchBronBron wants to merge 4 commits into
Open
TwitchBronBron wants to merge 4 commits into
TwitchBronBron wants to merge 4 commits into
Conversation
TwitchBronBron
force-pushed
the
ast-to-string
branch
from
August 15, 2025 12:12
22ecef7 to
2ac1c0f
Compare
Every AST node now has a `toSourceNode(state)` method and a `toString()` method that
rebuild the node's source code, including all leading trivia (whitespace, comments,
newlines, colons). For an unmodified AST, `ast.toString()` produces the exact text that
was parsed.
- Add `tokenToSourceNodeWithTrivia`, `nodeToSourceNode`, `nodesToSourceNode`, and
`statementsToSourceNode` helpers to `TranspileState`
- Store tokens the parser previously discarded: commas (call args, function params,
array literals, indexes, dim, callfunc, inline interfaces, typed function types),
template string `${`/`}` tokens, the `.` in `obj.[index]`, and the eof token on the root body
- Colons that are part of the AST (AA members, ternary, labels) are removed from the
leading trivia of the next token so they aren't written twice
- When there are syntax errors, tokens skipped by the parser are moved into the leading
trivia of the next token so the AST still represents the full source
- Plugin-created nodes/tokens (no location, no trivia) get sensible default spacing
- Fix swapped left/right parens in TypedFunctionTypeExpression
- Fix AAMemberExpression and AnnotationExpression clones dropping the comma and call args
- Fix `exitwhile` losing its leading trivia
- Fix infinite loop in `consumeUntil` when `#error` is the last line of a file
- Fix lexer including unexpected characters in the text of the next token
- Fix duplicated end tokens when `end sub`/`end class` is missing at the end of the file
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TwitchBronBron
force-pushed
the
ast-to-string
branch
from
September 26, 2026 19:44
5015d18 to
533459a
Compare
TwitchBronBron
marked this pull request as ready for review
September 26, 2026 19:45
Add tests for the less common paths of every AST output method (`toSourceNode`/`toString`, `transpile`, and `getTypedef`): nodes that are missing optional parts (as created by plugins or by the parser during error recovery), and output methods that are normally bypassed by their parent. Every line and branch added for `toString()` is now covered. Fixes found along the way: - `end [object Object]` was written for functions that have a `sub`/`function` token but no end token (transpile and typedef) - `TranspileState.sourceNode()` crashed when given an undefined locatable - `ExitStatement` crashed when transpiled without a loop type token - Plugin-created `exit`/`continue` loop type tokens were not separated by a space - `InterfaceMethodStatement.leadingTrivia` crashed when the function type token was missing - `createToken(TokenKind.At)` produced `at` instead of `@` - Plugin-created functions with no `sub`/`function` token got a leading space before the name Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Typedefs dropped comments that were above an annotation (they belong to the annotation's `@` token, not the statement), and wrote comments that were between an annotation and its statement above the annotation. Add `BrsTranspileState.getTypedefLeadingCommentsAndAnnotations` to write both in source order, and use it for every typedef that writes comments - `MethodStatement.leadingTrivia` skipped the modifiers (i.e. `public`, `override`), so comments above them were dropped from typedefs and from the transpiled method function Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The lexer reported characters it doesn't recognize (i.e. a stray `|` or `%`) and then dropped them, so `ast.toString()` couldn't reproduce them. They are now kept as `TokenKind.UnexpectedCharacter` tokens in the leading trivia of the next token. They are never added to the token list, so the parser, diagnostics, and transpile output are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds the ability to call
.toString()on any BrightScript/BrighterScript AST node and get the source code it represents, including all leading trivia (whitespace, comments, newlines, colons). For an unmodified AST,ast.toString()returns exactly the text that was parsed.This PR has been rebuilt on top of
v1, which already tracks trivia on every token, so most of the original v0 work (lexer trivia,leadingWhitespace) is no longer needed.How it works
AstNodehas a new abstracttoSourceNode(state: TranspileState): SourceNode, implemented by every statement and expression.toString()calls it with a blankTranspileState. Because it returns aSourceNode, the output carries source-map positions.TranspileStatehelpers:tokenToSourceNodeWithTrivianodeToSourceNode(also writes any annotations)nodesToSourceNode(writes a list of nodes with their separators)statementsToSourceNodeParser changes
Some tokens were consumed by the parser but never stored, so they are now kept on the AST:
tokens.commasonCallExpression,FunctionExpression,ArrayLiteralExpression,IndexedGetExpression,IndexedSetStatement,DimStatement,CallfuncExpression,InlineInterfaceExpression,TypedFunctionTypeExpression, andInterfaceMethodStatement.commas[i]is the comma after itemi.tokens.expressionBeginsandtokens.expressionEnds(the${and}tokens) on template strings.tokens.dotfor theobj.[index]syntax onIndexedGetExpressionandIndexedSetStatement.tokens.eofon the rootBody. It holds everything after the last statement.key: value, ternarya ? b : c,label:) are removed from the next token's leading trivia, so they aren't written twice.Plugin-created nodes
Nodes and tokens created by plugins usually have no location and no trivia. Written as-is, they would run into their neighbors (
x=5+3, statements on the same line). So when a token or node has no location and no leading trivia,toString()adds sensible defaults:=andas,between argumentsTokens and nodes that came from the parser are never changed.
Bugs fixed along the way
TypedFunctionTypeExpressionhad its left and right parens swapped.AAMemberExpressiondropped its comma, and cloning anAnnotationExpressiondropped its call arguments.exitwhilelost its leading trivia when it was split intoexit+while.#errorwas the last line of a file with no trailing newline."%\n".|) are now kept asTokenKind.UnexpectedCharacterleading trivia instead of being dropped, so they round-trip. They are never added to the token list, so parsing, diagnostics, and transpile output are unchanged.end suborend classwas missing at the end of a file, the previous token was stored a second time as the end token.createToken(TokenKind.ForEach),createToken(TokenKind.ExitWhile), andcreateToken(TokenKind.At)now default tofor each,exit while, and@.sub/functiontoken but no end token was transpiled (and typedef'd) asend [object Object].TranspileState.sourceNode()crashed when given an undefined locatable.ExitStatement.transpilecrashed without a loop type token, and plugin-createdexit/continueloop types were not separated by a space.InterfaceMethodStatement.leadingTriviacrashed when the function type token was missing.@token trivia), and moved comments between an annotation and its statement above the annotation. A newBrsTranspileState.getTypedefLeadingCommentsAndAnnotationswrites both in source order for every typedef.MethodStatement.leadingTriviaskipped thepublic/overridemodifiers, so comments above those methods were dropped from typedefs and from the transpiled method function.Verification
AstOutput.spec.tsalso covers previously untested paths in the existingtranspileandgetTypedefmethods, includingSOURCE_NAMESPACE_NAMEandSOURCE_NAMESPACE_ROOT_NAME, typedef leading comments and annotations, and nodes missing optional parts. The only uncovered branches left in those methods are defensive fallbacks that can't be reached (for examplegetLeadingComments(...) ?? [], since it always returns an array).ast.toString() === input, and every unique input matches exactly..brs/.bsfiles (rooibos, the Roku SDK samples, and a large production app) all round-trip exactly.public/overridemethods.Known limitations
trackLocations: true). WithtrackLocations: false, every token looks like it was created by a plugin, so default spacing may be added.trackLocations: falseis only used in tests.🤖 Generated with Claude Code