fix: decode BOM-less UTF-16 logs in the shared encoding fallback - #186
Conversation
The utf-16 codec needs a BOM and the utf-8 attempt succeeds on NUL bytes, so a BOM-less utf-16-le/be log was decoded as mojibake and yielded zero usages. Sniff the first two bytes and try the matching codec first. Addresses yunaremaia#185.
|
Merged — thanks. This was the right scope for a first PR: 21 lines, CI green, and both new tests fail on On the latin-1 half, my preference, with the reasoning so you can push back if you read the code differently: Do not keep
That second half is the part I would want settled before touching this again, and it is what #185 describes as consequence one: a genuinely binary log currently decodes as latin-1, every line fails One thing worth keeping in mind while you do it: And thank you for stating the tooling in the PR description. That is the right way to do it and it makes the change much easier to review and trust. |
|
Verified on the merged result, on top of what I noted on the latin-1 direction. The scope you picked was correct and the sniff holds up under the cases I could think of. Recording what I actually checked so the next person does not have to re-derive it. The two-byte sniff. Ran
The interesting question was whether a legitimate UTF-8 log can mis-trigger it. It cannot: the trigger requires Operator precedence. The line is correct, and I confirmed it by AST rather than by eye — the RHS node is an
Tests: red/green proof. I checked out the pre-PR parent and put only the PR's test file on top of it. Both new cases fail for exactly the reason in the issue, then pass with the change:
The parametrization genuinely covers both variants — Suite. 215 passed locally, 1 failed: One edge, deliberately not a blocker. A UTF-32 file (LE or BE) starts The CHANGELOG entry is accurate on every checkable claim: correct codecs named, correct issue number, right section, and the silent-zero-usage consequence is exactly what the red test demonstrates. On the second half of #185, the point that mattered to me was not adding a fourth codec but removing the escape hatch that makes the loop unconditionally succeed. Worth deciding deliberately when you get there — thanks again for keeping this one tight and for flagging the remaining scope rather than folding it in. |
|
I'd like to take the latin-1 half of #185. While checking it I found a crash. A clean latin-1 log with an accent, like "café", makes _parse_file raise UnicodeError: UTF-16 stream does not start with BOM. The utf-16 step raises UnicodeError when there is no BOM, and the loop only catches UnicodeDecodeError, so it never reaches latin-1 for this file. A directory scan with that file between two valid ones aborts the whole run. I ran the same file on the commit before #184 and it crashed there too, with a UnicodeDecodeError. My plan is to try utf-16 only when the file starts with a BOM, then utf-8. If both fail, the file is read as utf-8 with errors="replace", and on my test file the JSON still parsed. The reader then says the read was lossy, so the parsers can log the warning. One question. Should --strict make a lossy file fail the run? Only the Cursor parser has --strict today. |
|
Opened a PR for the latin-1 half: #193. |
Addresses #185 (the BOM-less UTF-16 part only).
Description
A UTF-16 log without a BOM was being read as UTF-8, because the NUL bytes are valid UTF-8. Every line then failed json.loads, so the file gave zero usages and no warning.
_read_lines_with_fallback now checks the first two bytes before anything else. If it looks like UTF-16 little-endian or big-endian, it uses that encoding first. Everything else goes through the same loop as before, so UTF-16 files with a BOM and normal UTF-8 files behave the same.
Type of change
Testing
Notes
I left the latin-1 and unreachable None part of #185 alone. I asked on #182 which way you prefer, and I'll do it once you say.
The check only catches this when the first character is ASCII, which is true for JSON logs.
AI use: written with Claude Sonnet 5.5. I ran and checked every command and test myself.