Skip to content
144 changes: 136 additions & 8 deletions docs/testing/BINARY_S3_STORAGE.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,20 @@ table may not be seen at first read, and a runtime change is ignored until resta
overrides through `Config.setProperty` do re-read it, which is how tests switch modes. Tests
that mock `Config` statically must call `AssetStorageFeature.reset()` themselves.

With the flag off, the S3 cleanup job queues (`binaryAssetCleanup`, `binaryFieldCleanup`) are not
registered.

### Enabling the flag is not rollback-safe

Enabling the flag is a one-way step for any content written while it is on. Check-in stores the
active binary only under a `.revisions/<uuid>/` key recorded in `contentlet_as_json`
(`storageKey`/`metadataStorageKey`); no legacy flat file is written, and the local copy may later
be evicted to S3. Neither a release without this code nor this release with the flag turned back
off reads those revision keys (the reader falls back to the legacy field folder), so affected
binaries resolve as stale or missing. Leaving the flag off, the default, changes nothing. Keeping
a legacy-path copy for rollback is not implemented; treat enabling the flag as requiring a
forward-only recovery plan.

## Storage layer behavior with the flag on

The storage chain (`ChainableStoragePersistenceAPI`) and its providers change as follows:
Expand All @@ -59,9 +73,6 @@ The storage chain (`ChainableStoragePersistenceAPI`) and its providers change as
per lookup. Only listings that need every key (`listObjectPaths`, `listObjectSnapshots`,
`deleteGroup` and static push) follow every page.

`SharedExtractedMetadata` caches byte-derived Tika extraction at
`extracted-metadata/<source-sha256>/<configuration-hash>.json`. Nothing calls it yet; metadata
generation starts using it in a later slice.

## Binary asset API

Expand Down Expand Up @@ -92,6 +103,105 @@ leave its `{inode}/{field}` directory. This applies with the flag on or off.

With the flag off, the API uses the existing filesystem/NFS paths and makes no remote calls.

## Content binaries and immutable revisions

With the flag on, new CMS binary writes use
`binary-assets/{a}/{b}/{inode}/{field}/.revisions/{revision-id}/{filename}`. The Binary field JSON
keeps its original `value` filename and adds `storageKey` (and `metadataStorageKey` for its
metadata). The field reference changes in the content transaction, so an uncommitted replacement
does not overwrite the prior object, and a rollback keeps the previous revision. Legacy binary
JSON and storage keys remain readable. Reconstructed files keep the stored revision path without
I/O. Old revisions are kept until whole-inode cleanup. A check-in that rolls back deletes the
revision and revision metadata it uploaded, through a rollback listener; the key carries a fresh
UUID, so no other version can reference it. If that delete fails it is logged and the object is
left behind. Metadata written under a different key during the rolled-back check-in, and
revisions abandoned by a rollback to a savepoint, are not reclaimed.

Custom metadata is copied to replacements, and metadata files and cache entries identify the exact
binary revision. Binary HTTP responses (`BinaryExporterServlet`) hold a cache lease while they
resolve, export and open the file, and release it before streaming, so a slow client does not
defer eviction. `FileAsset.getInputStream` and `Contentlet.getBinaryStream` also hold it only until
the stream is open, because a lease must be released on the thread that took it and a caller may
read or close the stream elsewhere. On a local disk an open file survives eviction, so this is
safe. Eviction on an NFS asset directory, where deleting an open file can break the read,
has not been validated.

## Metadata from evicted originals

With the flag on, `FileStorageAPI` restores the exact binary only when metadata generation needs
its bytes, and holds a cache lease through basic inspection, hashing and Tika parsing. Restore
failures propagate, and so do shared-extraction storage failures, instead of producing a
successful partial metadata result.

Byte-derived Tika extraction is shared through `SharedExtractedMetadata`, cached at
`extracted-metadata/<source-sha256>/<configuration-hash>.json`. The configuration hash covers the
parser bundle version, the binary metadata schema version and the extracted-text limit. Filenames,
local paths, fallback titles and modification times are added per use afterwards, and custom
attributes and focal points stay in their content-owned snapshots. Unknown parser versions and
flag-off mode extract directly. Empty or failed extractions are not published. Shared extraction
records are not reclaimed yet.

JSON metadata hydration reads the linked image's owner key, not the parent content's binary key.
Complete stored metadata avoids normalizing or downloading the original. Missing metadata can be
regenerated from a cold legacy or revision path.

## Durable deletion

Whole-inode deletion, deletion of one language of multilingual content, and old-version
maintenance record one `binaryAssetCleanup` job per deleted inode in the same database transaction
as the content deletion. Main leaves the files of a deleted language on disk (#9146); with the
flag on, that path now records cleanup jobs and, when `BACKUP_DELETED_CONTENTLETS_TO_DISK` is on,
recovery archives like the other deletion paths.

The job carries the exact binary and metadata paths stored for the inode when it was recorded, so
recording it lists the inode's objects inside the deletion transaction, and a storage listing
failure aborts the deletion. The worker refuses while a content version with that inode exists,
then deletes the recorded metadata before the recorded source objects, so a failure leaves the
sources available for a retry. It never deletes by prefix: an inode can be re-created after the
deletion (push publishing keeps the sender's inodes), and the revisions it uploads survive. Objects
uploaded under the deleted inode by a transaction that overlapped the deletion are therefore not
reclaimed. The inode's completed renditions and legacy image cache are still removed whole, since
they regenerate on demand. S3 failures use the job queue's retry policy; after retries are
exhausted the job stays failed and can be retried through the job management API. Direct
submissions through the public job endpoint are rejected. With the flag off, workers do not touch
storage and pending jobs are not silently completed.

## Binary field trash

With the flag on, deleting a binary field records a `binaryFieldCleanup` request in the same
transaction as the field deletion. Requests cover historical and working versions up to the
deletion time; values from a later field with the same name are kept.

Cleanup runs as a queued job, never inside the caller's request. It handles one content row per
transaction: it locks that row, rechecks that it was not edited after the field was removed,
uploads and verifies a recovery ZIP, clears the old field reference, records the exact cleanup
inventory, and advances the job's saved cursor, all in the row's own commit. Only that row is locked
while its archive uploads, and a retry resumes after the last committed row. A separate step checks that
the archived files are no longer referenced and that the ZIP is still available before deleting
metadata, originals and renditions. It never deletes a whole field prefix, so uploads made after
the inventory was captured survive a retry. Direct submissions through the public job endpoint are
rejected; only field deletion and `ContentletAPI.cleanField` create this work. With the flag off, scheduling and local
trash behave as before.

## Deleted-content recovery archives

When both the flag and the existing `BACKUP_DELETED_CONTENTLETS_TO_DISK` option are on, deletion
writes a verified S3 recovery ZIP before removing content. Archives live in the
`deleted-content-backups` group at `<identifier>/<inode>/<uuid>.zip`. Full destruction and
all-version deletion archive each version, deleting one language archives each version in that
language, and single-version deletion archives only that version. A failed backup aborts the
deletion. Binary cleanup never deletes recovery archives. Field trash ZIPs use the same group and
layout.

Each ZIP contains `contentlet.json` (the row's complete typed field data), `contentlet.xml`, and
`assets/` entries under the original binary and metadata paths, including binary fields whose
definitions were removed. Operators can download the ZIPs from S3 for manual recovery; internal
callers can use `ContentletBackupStorage.list(identifier)` and `open(key)`. There is no automatic
database restore, and archives are kept until removed or expired by the bucket's lifecycle policy.

With the flag on, the `deleteAllVersionsandBackup` interceptor, previously a no-op, calls its
implementation and the all-version deletion hooks. With the flag off it stays a no-op.

## Local cache eviction

With the flag on, the local asset directory is a cache that `BinaryCacheEvictionJob` can trim.
Expand Down Expand Up @@ -171,8 +281,10 @@ an S3-compatible custom endpoint needs a key and secret.
## Running the checks

The S3 checks use the real filesystem provider, chain and AWS adapter against a disposable
MinIO bucket. `BinaryS3StorageTest` is skipped unless `s3.test.endpoint` is set; the other
tests always run. The credentials below are disposable local test values.
MinIO bucket, and the transaction checks use a disposable PostgreSQL database; each test creates
and removes its own schema. `BinaryS3StorageTest` is skipped unless `s3.test.endpoint` is set, and
the PostgreSQL cases are skipped unless `s3.test.jdbc` is set. The credentials below are disposable
local test values.

```sh
docker run -d --rm --name binary-s3-test \
Expand All @@ -181,11 +293,27 @@ docker run -d --rm --name binary-s3-test \
-e MINIO_ROOT_PASSWORD=binary-storage-test \
minio/minio:latest server /data

docker run -d --rm --name binary-cleanup-postgres-test \
-p 127.0.0.1:19003:5432 \
-e POSTGRES_USER=binary-storage-test \
-e POSTGRES_PASSWORD=binary-storage-test \
-e POSTGRES_DB=binary_storage_test postgres:16-alpine

./mvnw test -pl :dotcms-core -Dmaven.build.cache.enabled=false \
-Dtest=AssetStorageFeatureTest,AssetStorageFeatureLatchTest,S3StorageConfigurationTest,NoWebIdentityCredentialsProviderChainTest,BinaryS3StorageTest,BinaryAssetReferenceTest,BinaryCacheEvictionJobTest,BinaryFileSystemStorageTest,BinaryAssetStorageAPIImplTest,MetadataLocalCacheTest \
-Ds3.test.endpoint=http://127.0.0.1:19002
-Dtest=AssetStorageFeatureTest,AssetStorageFeatureLatchTest,S3StorageConfigurationTest,NoWebIdentityCredentialsProviderChainTest,BinaryS3StorageTest,BinaryAssetReferenceTest,BinaryCacheEvictionJobTest,BinaryFileSystemStorageTest,BinaryAssetStorageAPIImplTest,MetadataLocalCacheTest,BinaryAssetCleanupTransactionTest,BinaryAssetCleanupProcessorTest,ContentletBackupStorageGateTest,BinaryFieldCleanupProcessorTest,AssetJobEventSerializationTest \
-Ds3.test.endpoint=http://127.0.0.1:19002 \
-Ds3.test.jdbc=jdbc:postgresql://127.0.0.1:19003/binary_storage_test

docker stop binary-s3-test
docker stop binary-s3-test binary-cleanup-postgres-test
```

The CMS integration checks are registered in `Junit5Suite1` and run against the full integration
stack:

```sh
./mvnw install -pl :dotcms-core --am -DskipTests -Ddocker.skip
./mvnw verify -pl :dotcms-integration -Dmaven.build.cache.enabled=false -Dcoreit.test.skip=false \
-Dit.test=BinaryAssetStorageIntegrationTest,ContentletBackupStorageTest,SharedAssetStorageIntegrationTest
```

CI does not yet provide the MinIO service, so `BinaryS3StorageTest` does not run there.
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,6 @@
import com.dotmarketing.exception.DotSecurityException;
import com.dotmarketing.portlets.categories.business.CategoryAPI;
import com.dotmarketing.portlets.categories.model.Category;
import com.dotmarketing.portlets.contentlet.business.BinaryFileFilter;
import com.dotmarketing.portlets.contentlet.business.ContentletAPI;
import com.dotmarketing.portlets.contentlet.business.HostAPI;
import com.dotmarketing.portlets.fileassets.business.FileAssetAPI;
Expand Down Expand Up @@ -83,8 +82,6 @@
*/
public class ContentletJsonAPIImpl implements ContentletJsonAPI {

private static final BinaryFileFilter binaryFileFilter = new BinaryFileFilter();

final IdentifierAPI identifierAPI;
final ContentTypeAPI contentTypeAPI;
final FileAssetAPI fileAssetAPI;
Expand Down Expand Up @@ -401,6 +398,15 @@ private boolean isNotMappable(final Field field) {
*/
private Optional<File> getBinary(final Field field, final String inode,
final FieldValue<?> storedValue) {
if (com.dotcms.storage.AssetStorageFeature.isEnabled()
&& storedValue instanceof com.dotcms.content.model.type.system.AbstractBinaryFieldType) {
final String key = ((com.dotcms.content.model.type.system.AbstractBinaryFieldType) storedValue).storageKey();
if (key != null) {
return Optional.of(com.dotcms.storage.binary.BinaryAssetReference.withMetadata(
com.dotcms.storage.binary.BinaryAssetReference.localFile(inode, field.variable(), key),
inode, field.variable(), ((com.dotcms.content.model.type.system.AbstractBinaryFieldType) storedValue).metadataStorageKey()));
}
}
// This validation is here to prevent an exception.
// Cause the json gets saved twice by internalCheckin and the first time it does it no inode is set yet

Expand All @@ -422,17 +428,31 @@ private Optional<File> getBinary(final Field field, final String inode,
final Object storedName = null != storedValue ? storedValue.value() : null;
if (storedName instanceof String && isSet((String) storedName)
&& !((String) storedName).contains("/") && !((String) storedName).contains("\\")) {
return Optional.of(new java.io.File(binaryFileFolder, (String) storedName));
final File file = new java.io.File(binaryFileFolder, (String) storedName);
if (com.dotcms.storage.AssetStorageFeature.isEnabled()
&& storedValue instanceof com.dotcms.content.model.type.system.AbstractBinaryFieldType) {
return Optional.of(com.dotcms.storage.binary.BinaryAssetReference.withMetadata(file, inode,
field.variable(), ((com.dotcms.content.model.type.system.AbstractBinaryFieldType) storedValue).metadataStorageKey()));
}
return Optional.of(file);
}

// Legacy json without a stored file name: fall back to listing the folder. No exists()
// pre-check — listFiles() returns null for a missing folder.
final java.io.File[] files = binaryFileFolder.listFiles(binaryFileFilter);
if (files != null && files.length > 0) {
return Optional.of(files[0]);
if (!com.dotcms.storage.AssetStorageFeature.isEnabled()) {
final File[] files = binaryFileFolder.listFiles(new com.dotmarketing.portlets.contentlet.business.BinaryFileFilter());
return files != null && files.length > 0 ? Optional.of(files[0]) : Optional.empty();
}

return Optional.empty();
// Legacy json without a stored filename resolves through binary storage.
try {
final File file = APILocator.getBinaryAssetStorageAPI()
.getBinaryFile(inode, field.variable());
return Optional.ofNullable(file);
} catch (final DotDataException e) {
Logger.debug(this, () -> String.format(
"Binary not found for inode '%s', field '%s': %s",
inode, field.variable(), e.getMessage()));
return Optional.empty();
}
}

/**
Expand Down Expand Up @@ -474,6 +494,14 @@ private Optional<FieldValue<?>> hydrateThenGetFieldValue(final Object value, fin
final Optional<FieldValueBuilder> fieldValueBuilder = field.fieldValue(value);
if (fieldValueBuilder.isPresent()) {
FieldValueBuilder builder = fieldValueBuilder.get();
if (com.dotcms.storage.AssetStorageFeature.isEnabled() && value instanceof File
&& builder instanceof com.dotcms.content.model.type.system.BinaryFieldType.Builder) {
((com.dotcms.content.model.type.system.BinaryFieldType.Builder) builder).storageKey(
com.dotcms.storage.binary.BinaryAssetReference.keyOf((File) value,
contentlet.getInode(), field.variable()))
.metadataStorageKey(com.dotcms.storage.binary.BinaryAssetReference.metadataKeyOf(
(File) value, contentlet.getInode(), field.variable()));
}
final List<Tuple2<HydrationDelegate,String>> delegateAndFields = getHydrationDelegatesFromAnnotations(builder.getClass());
for (Tuple2<HydrationDelegate,String> delegateAndField : delegateAndFields) {
final HydrationDelegate delegate = delegateAndField._1();
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -173,10 +173,17 @@ public int deleteOldContent() {
dc.setSQL(String.format(DELETE_TAG_INODES, inodes));
dc.loadResult(conn);

if (com.dotcms.storage.AssetStorageFeature.isEnabled() && CLEAN_DEAD_INODE_FROM_FS) {
for (final String inode : inodeList) {
com.dotcms.storage.binary.BinaryAssetCleanupProcessor.enqueue(inode);
}
}
conn.commit();
conn.setAutoCommit(true);

deleteFromAssetsDir(inodeList);
if (!com.dotcms.storage.AssetStorageFeature.isEnabled()) {
deleteFromAssetsDir(inodeList);
}

inodeList.clear();
if (isInterrupted()) {
Expand Down
Loading
Loading