Skip to content

+ MOBI 7 books are created from one HTML document, and real books' text is read - #403

Merged
Hawkynt merged 3 commits into
mainfrom
feat/mobi-azw-support
Sep 30, 2026
Merged

Hawkynt merged 3 commits into
mainfrom
feat/mobi-azw-support

Conversation

@Hawkynt

@Hawkynt Hawkynt commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

What changed

  • Create: one UTF-8 HTML document becomes a MOBI 7 book: record 0 (PalmDOC header, 232-byte MOBI header, EXTH, full name), 4096-byte stored or PalmDOC text records, and an end-of-file record. The options are title, author, publisher, description, ISBN, subject and language. Encryption is refused.
  • Read: stored and PalmDOC text is decoded into book.html, with the text records' trailing entries stripped first (extra record data flags: backward-encoded sizes, multibyte overlap). HUFF/CDIC and DRM text stay raw records only.
  • Offsets: the reader looked for the text encoding, locale and EXTH flags 16 bytes too far (MobileRead offsets are from record 0). On real books it reported encoding 0xFFFFFFFF and never found EXTH. The first writer on this branch put every field from 0x50 on in the same wrong places, so its self round trip passed. Both are fixed, and the full name is read too.

Verification (not through our own reader)

  • A calibre 4.17 book, Project Gutenberg #1065 (public domain, embedded with provenance): book.html equals KindleUnpack 0.4.1's getRawML() byte for byte, and title, author, encoding and extra-data flags are read from their offsets.
  • Our writer's output, both methods, read by KindleUnpack: title, author, codec and the text match. The test runs where CWB_KINDLEUNPACK_PATH points at the python mobi package and is skipped elsewhere. It passed locally.
  • Record-0 fields are asserted at the MobileRead offsets for 0, 1, 4096, 4097 and 12288 bytes of text.
  • Trailing-entry, PalmDOC-decoder (malformed input, expansion cap) and encoder round-trip boundary cases.
  • dotnet test --filter Mobi|SupportMatrix|BundledDescriptor|CapabilityDocumentation|RoundTripsItsOwnOutput: 195 passed.

Sourcing: rung 4, the MobileRead MOBI, PDB and PalmDOC pages. No implementation code was consulted. KindleUnpack was used only as a black-box oracle.

Not verified: an actual Kindle device or Kindle Previewer. No FLIS, FCIS or index records are written.

@Hawkynt Hawkynt changed the title Decode stored and PalmDOC-compressed MOBI text Create MOBI 7 books and decode stored/PalmDOC text Sep 29, 2026
@Hawkynt
Hawkynt force-pushed the feat/mobi-azw-support branch from 3b702e2 to 465a4c1 Compare September 30, 2026 12:25
@Hawkynt Hawkynt changed the title Create MOBI 7 books and decode stored/PalmDOC text + MOBI 7 books are created from one HTML document, and real books' text is read Sep 30, 2026
@Hawkynt
Hawkynt force-pushed the feat/mobi-azw-support branch from 5a97ebf to e871872 Compare September 30, 2026 14:54
…xt is read

Creation: one UTF-8 HTML input becomes record 0 (PalmDOC header,
232-byte MOBI header, EXTH with title/author/publisher/description/
ISBN/subject, full name), 4096-byte stored or PalmDOC text records, and
an end-of-file record. Encryption is refused.

Reading: stored and PalmDOC text is decoded into book.html. The text
records' trailing entries (extra record data flags: backward-encoded
sizes, multibyte overlap) are stripped first. HUFF/CDIC and DRM text stay
available as raw records only.

The reader had looked for the text encoding, locale and EXTH flags
sixteen bytes past where they are, as if the MobileRead offsets were
counted from the MOBI magic rather than from record 0. On real books it
reported encoding 0xFFFFFFFF and never found EXTH. The first writer on
this branch put every field from 0x50 on in the same wrong places, so
its round trip passed. Both now use the record-0 offsets, and the full
name is read as well.

Rung 4 (MobileRead MOBI, PDB and PalmDOC pages); no implementation code
consulted. Verified two ways, neither through our own reader:
- a calibre 4.17 book (Gutenberg #1065, embedded): book.html equals
  KindleUnpack 0.4.1's getRawML() byte for byte (SHA-1 recorded), and
  title, author and encoding are read from their offsets;
- our writer's output, stored and PalmDOC, read by KindleUnpack: title,
  author, codec and the text match. The test runs where
  CWB_KINDLEUNPACK_PATH points at the python 'mobi' package. The
  record-0 fields are also asserted directly at the MobileRead offsets
  for 0, 1, 4096, 4097 and 12288 bytes of text.

Trailing-entry, PalmDOC-decoder (malformed input, expansion cap) and
encoder round-trip cases cover the boundaries. The package README's
corrupted copy from this branch was replaced with main's plus the
Mobi row.
@Hawkynt
Hawkynt force-pushed the feat/mobi-azw-support branch from 79fb252 to 5e1a1b4 Compare September 30, 2026 16:07
@Hawkynt
Hawkynt merged commit 2699deb into main Sep 30, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant