Skip to content

feat(storage): add feature-gated S3 storage layer for shared asset storage - #37770

Open
swicken wants to merge 6 commits into
mainfrom
s3-stack/1a-storage-layer
Open

swicken wants to merge 6 commits into
mainfrom
s3-stack/1a-storage-layer

Conversation

@swicken

@swicken swicken commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

S3 asset storage, part 1 of 7. Stacked PRs, review bottom up. Each one builds on the one below.
1a storage layer #37770 · 1b binary asset API #37771 · 2 content #37772 · 3 recovery #37773 · 4 publishing #37774 · 5 temporary uploads and WebDAV #37775 · 6 rendering #37776
Everything is behind FEATURE_FLAG_S3_ASSET_STORAGE, off by default. With the flag off, behavior matches main.

Refs #37868

Proposed Changes

This is the foundation for storing binary assets durably in S3, with the local asset directory as a cache. It adds the flag and the storage-layer behavior the later PRs build on. Content binaries are not stored through S3 until 2; with the flag on, the only paths this PR changes for running code are metadata reads and writes through FileStorageAPI (storage failures now propagate instead of reading as missing) and static push publishing.

  • Feature flag. AssetStorageFeature reads FEATURE_FLAG_S3_ASSET_STORAGE once per process and keeps the value, because several storage objects capture the mode when they are built. It is set through the environment or dotmarketing-config.properties and needs a restart; a runtime change is ignored. Config.setProperty resets the cached value so tests can switch modes. The first read logs at INFO only when the flag is on.
  • Storage chain behavior with the flag on (ChainableStoragePersistenceAPI and the filesystem, database and S3 providers): writes publish to durable providers before the new local copy becomes visible, reads restore missing local copies from S3 without caching the miss, deletes keep the local copy until every durable provider accepts, and storage or database failures propagate instead of reading as a missing object. Restores take the same per-key lock as writes and deletes. A zero-length or truncated local metadata file is replaced from S3 when S3 holds a readable copy; otherwise the read fails and the file is kept, so the metadata is never treated as absent and regenerated. Providers gain prefix listing, durable-copy verification and non-overwriting backfill; S3 existence checks list at most one key.
  • S3 configuration. Default AWS credential chain when no keys are set (the STS module is packaged for role-based credentials), region and endpoint validation, and an optional storage.file-metadata.s3.namespace to separate installations that share a bucket.
  • Building blocks for later PRs: S3ContentAddressedStorage (one immutable blob per set of bytes, used from 1b) and SharedExtractedMetadata (shared Tika extraction, used from 2).
  • Static push publishing to S3 (AWSS3Storage) signs with SigV4, uses the bucket's own region (looked up by the SDK when none is configured), and lists every page of objects, but only when the flag is on.
  • docs/testing/BINARY_S3_STORAGE.md describes the current behavior and how to run the checks. Later PRs add their own sections.

Behavior with the flag off

Unchanged from main, with one documented difference in AWS credential resolution. Every new branch in the providers is gated on the flag, and the new configuration keys are ignored.

The STS module is now on the classpath. Where dotCMS itself falls back to the AWS default chain (static push publishing, its endpoint validation and the S3 metadata provider), it uses NoWebIdentityCredentialsProviderChain while the flag is off: the SDK's default chain without the web-identity step, which always failed over on main. So a pod with a web-identity token keeps the identity it had on main. Two cases still see the module: a plugin that builds its own DefaultAWSCredentialsProviderChain, and an AWS profile that uses role_arn, which used to fail over to the next provider and now assumes the role.

Review fixes

The last commit on this branch (fix(storage): keep flag-off AWS credential resolution and harden the flag-on storage layer) addresses a full review of this PR:

  • Flag-off credential resolution and the flag-off INFO line, as described above.
  • The chain's object restore now takes the per-key lock, so a restore can no longer overwrite a newer write or bring back a deleted object.
  • An unreadable local metadata file no longer fails every read when S3 has a good copy (the new UnreadableStoredObjectException).
  • Storage.listFirstObject replaces whole-group listings in existsGroup, existsObject and hasDurableCopy.
  • With the flag on, static push clients with no region and no custom endpoint use the global endpoint and the SDK's region lookup instead of failing with "Unable to find a region", and the credential-chain client is no longer pinned to us-west-2.
  • The S3 pushFile starts its upload inside its lock, the filesystem prefix listing resolves its prefix the way writes do, and unreachable flag checks are removed.
  • The doc now states that both asset-blobs/ and extracted-metadata/ are shared across namespaces.

Deliberately not in this PR

The binary asset API and eviction (1b). Anything that routes content, metadata or other subsystems through S3 (2 to 6). The job-processor registration filter moves to 2, where the first S3 job processors arrive.

Checklist

  • Tests: the 25 unit tests in the doc's run command pass after the review fixes, including the MinIO cases against a disposable bucket, and the new NoWebIdentityCredentialsProviderChainTest. The MinIO tests skip in CI because CI provides no S3 endpoint (-Ds3.test.endpoint).
  • Translations: not applicable.
  • Security Implications Contemplated: S3 credentials come from existing config or the AWS default chain and are never logged. The namespace separates application keys only; it is not an access boundary, and the doc says so.

Additional Info

This is the only PR in the stack that gets the full PR CI run, because that workflow only runs for PRs based on main. Rollback: merging with the flag off is safe to roll back. With the flag on and a namespace configured, metadata keys move under asset-namespaces/..., which an older release does not read; the one-way content change starts in 2.

@claude

claude Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @fabrizzio-dotCMS's task in 3m 11s —— View job


Code Review — S3 storage layer (PR #37770)

I traced every new branch in the storage providers and confirmed each is gated on AssetStorageFeature.isEnabled(), so flag-off behavior matches main (with the one documented AWS-credential difference, which NoWebIdentityCredentialsProviderChain handles correctly). The flag-on paths — durable-first writes under a per-key Striped lock, refill-on-read under that same lock, UnreadableStoredObjectException handling, single-key existence listing (listFirstObject), and conditional/multipart uploads with abort-on-failure — hold up. I found no new provable P0/P1 regression beyond what prior reviewers already filed.

Below is my recheck of the open findings plus a couple of smaller observations.

New Issues

  • 🟡 Medium: ChainableStoragePersistenceAPI.java:490 — flag-on pullObject restore now performs a synchronous durable write inside the read path (cacheStorage.pushObject(...) for every missing provider, including the DB). On main, refill was best-effort async (pushObjectAsync, failures swallowed). Now, if a refill write to a durable provider fails, the read throws DotRuntimeException and the caller sees a read failure for an object that genuinely exists downstream. This is consistent with the PR's "failures propagate" stance and is gated, so flagging as non-blocking — but it widens the blast radius of a transient durable-write error into read failures. Assumption: the DB/S3 providers can be in the missing list on a read. What to verify: that a failed restore-write is an acceptable read outcome vs. logging + continuing to serve the readable copy.

Existing

  • 🟠 High: ChainableStoragePersistenceAPI.java:259 / :292 — stale local copy after a durable write succeeds but the local write then fails (fabrizzio's finding, unaddressed). S3 holds the new value, the local file keeps the old one, and this node serves the old value indefinitely with no error; eviction (feat(storage): add BinaryAssetStorageAPI and local cache eviction for S3 asset storage #37771) can't distinguish it from a failed local-only rendition. Still the most important item in the PR. The suggested fix (best-effort deleteObjectAndReferences on the local provider before rethrowing, mirrored in pushObject) still applies. Fix this →
  • 🟡 Medium: AmazonS3StoragePersistenceAPIImpl.java:419 — pushFile holds a stripe of the process-wide shared IdentifierStripedLock across the whole upload. Harmless for small metadata JSON here, but from feat(storage): store push-publishing bundles durably in S3 asset storage #37774 large bundle .tar.gz uploads route through this method and will pin a shared stripe for the full transfer, so unrelated content saves hashing to the same stripe fail with DotConcurrentException. fabrizzio's suggestion (drop the lock on the flag-on branch, since every flag-on caller already serializes per key; document that pushFile doesn't order concurrent same-key writes) still applies.
  • 🟡 Medium: SharedExtractedMetadata.java:34 — null extractorVersion collapses distinct parser bundles into one cache key ("tika:null:..."), violating the class's own "unknown implementations must not share extraction caches" contract. TikaUtils.extractorVersion() (TikaUtils.java:671) returns null when OSGi/Tika is uninitialized. Unreachable in this PR (no in-patch caller), but it is wired in feat(storage): route content binaries, metadata and cleanup through S3 asset storage #37772, so worth guarding now. Fix this →
  • 🟡 Medium: AmazonS3StoragePersistenceAPIImpl.java:820 / :856 — backfillFile/backfillObject treat only 412 as "someone else wrote first", but uploadFileIfMatch and S3ContentAddressedStorage.preconditionFailure also accept 409 (ConditionalRequestConflict). A 409 here throws instead of falling through to the hasDurableCopy re-check, failing the batch. Treat 409 like 412 (or share one preconditionFailure helper) for consistency across the three paths.
  • 🟡 Medium: FileSystemStoragePersistenceAPIImpl.java:437 — flag-on pullObject treats only Jackson/EOFException as unreadable. With CONTENT_METADATA_COMPRESSOR=gzip, a corrupt header throws ZipException; with bzip2, a zero-length file throws a format IOException. Both become DotDataException, so the chain fails instead of replacing the file from S3. Default is none, so only compressed setups are affected.
  • 🟡 Medium: S3ContentAddressedStorage.java:108 — store downloads the full blob twice to verify (blobMatches after upload + matches at the end), and matches re-hashes the local snapshot. From feat(storage): route content binaries, metadata and cleanup through S3 asset storage #37772 this runs inside the check-in transaction with the contentlet row locked, so a 1 GB binary keeps the transaction open for the upload plus 2 GB of reads. Compare blob bytes only when the blob pre-existed, and end with the cheap reference-header check.
  • 🟡 Medium: AmazonS3StoragePersistenceAPIImpl.java:102 — lock-key case mismatch: the chain's lock key (groupName + "/" + path) is case-sensitive and S3 keys keep case, but the filesystem provider lowercases every path (normalizePath). Two keys differing only in case → two S3 objects / two locks but one local file, so the lock's ordering guarantee doesn't hold for that pair. fabrizzio notes it may be unreachable for the metadata chain (keys from inodes/field names); if so, stating that here (or normalizing the lock key) would close it.
  • 🟡 Medium: AssetStorageFeature.java:14 — FEATURE_FLAG_S3_ASSET_STORAGE should be declared in com.dotcms.featureflag.FeatureFlagName (with FLAG referencing it) like IndexConfigHelper.FLAG_KEY → FeatureFlagName.FEATURE_FLAG_OPEN_SEARCH_PHASE, so the flag is discoverable where the others live.
  • 🟡 Medium: BinaryS3StorageTest.java:30 — the only test touching a real S3 API is @EnabledIfSystemProperty("s3.test.endpoint") and CI never sets it, so If-None-Match/If-Match, listObjects pagination, and real 404/409/412 mapping never run in CI. Adding MinIO to the docker-maven-plugin imagesMap in parent/pom.xml (next to database/opensearch) would cover S3ContentAddressedStorage, the multipart branch of uploadFileIfAbsent (>32 MiB), and the stale-local-copy case.
  • 🟡 Medium: AmazonS3StoragePersistenceAPIImpl.java:794 — minor efficiency: hasDurableCopy reads the whole local file for its MD5 before confirming the key exists and the size matches; moving the digest after the listFirstObject/size check avoids a full read when the answer is already false.
  • 🟡 Medium: ChainableStoragePersistenceAPI.java:145 — the assetLocks Striped is per-instance and lock() waits forever uninterruptibly while doing real S3/NFS I/O. Making it static (process-wide) and using tryLock with a generous timeout would survive StoragePersistenceProvider.forceInitialize() rebuilds and keep a stuck holder from parking every thread on that stripe.
  • 🟡 Medium: Config.java:800 — stray ponytail: token at the start of the comment; drop it. ChainableStoragePersistenceAPI.java:38 — inline fully-qualified com.google.common.util.concurrent.Striped / java.util.concurrent.locks.Lock should be imported.

Resolved

  • ✅ ChainableStoragePersistenceAPI.java:471 — object restore now takes the per-key lock, so a restore can no longer overwrite a newer write or revive a deleted object.
  • ✅ FileSystemStoragePersistenceAPIImpl.java:439 — unreadable local metadata no longer fails every read; UnreadableStoredObjectException lets the chain replace it from S3 when S3 holds a readable copy.
  • ✅ AmazonS3StoragePersistenceAPIImpl.java:236,:268 / AWSS3Storage.java:260 — listFirstObject replaces whole-group listings in existsGroup/existsObject/hasDurableCopy; S3 existence checks list at most one key.
  • ✅ AWSS3Storage.java:50,:97 — flag-on clients with no region and no custom endpoint use the global endpoint + SDK region lookup instead of failing with "Unable to find a region"; the credential-chain client is no longer pinned to us-west-2, and the SigV4 signer override is dropped only when the flag is on.
  • ✅ NoWebIdentityCredentialsProviderChain.java — flag-off credential resolution preserves the pre-STS behavior; covered by NoWebIdentityCredentialsProviderChainTest.

Verdict: No new blocking bugs. The one pre-existing 🟠 High (stale local copy after a durable write succeeds) is the item I'd most want resolved before merge; the remaining 🟡 items are non-blocking but several (409 handling, compressor exceptions, MinIO CI coverage) are cheap and would meaningfully harden the flag-on path.

· s3-stack/1a-storage-layer

@swicken
swicken marked this pull request as ready for review October 2, 2026 18:24
@nollymar nollymar added the PR : dotbot review Trigger dotbot AI code review and the post-merge QA test plan label Oct 6, 2026
try {
Files.copy(source.toPath(), snapshot, StandardCopyOption.REPLACE_EXISTING);
final String hash = S3ContentAddressedStorage.hash(snapshot.toFile());
final String configuration = org.apache.commons.codec.digest.DigestUtils.sha256Hex(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 [P2] SharedExtractedMetadata.java:35 guard null extractorVersion before building shared cache key

Current code:

final String configuration = org.apache.commons.codec.digest.DigestUtils.sha256Hex(
        "tika:" + extractorVersion + ":schema:" + schemaVersion + ":text-limit:" + textLimit);

Problem: TikaUtils.extractorVersion() (TikaUtils.java:672) returns null when OSGi/Tika is uninitialized, so the key becomes "tika:null:..." and unknown parser bundles share one cached entry, violating the class's own "unknown implementations must not share extraction caches" contract. Dead until PR 2 wires it.

Fix:

if (extractorVersion == null || extractorVersion.isBlank()) return extract.apply(source);

Assumption: a PR-2 caller passes TikaUtils.extractorVersion() through unchanged. What to verify: no caller can reach get() with a null version and expect a correct per-parser cache.

…orage

First slice of the S3 asset storage work. Adds the FEATURE_FLAG_S3_ASSET_STORAGE
startup flag (read once per process), chain-compatible S3, filesystem and database
persistence operations, content-addressed S3 blob storage, shared extracted
metadata, STS credential support, and SigV4/region/pagination handling for static
push publishing. Everything is inert while the flag is off, which is the default.
…nceAPI and route remote storage through the provider
…flag-on storage layer

With the flag off, every place dotCMS falls back to the AWS default credential chain (static
push publishing, endpoint validation and the S3 metadata provider) now uses
NoWebIdentityCredentialsProviderChain, the SDK default chain without its web identity step.
Before the STS module was packaged that step always failed over, so a pod with a web identity
token keeps the identity it used before. With the flag on the standard chain is unchanged.

AssetStorageFeature logs at INFO only when the flag is on; with it off the first read logs at
debug, so a flag-off node writes no new INFO line.

The chain's pullObject now restores missing local copies under the same per-key lock as
writes and deletes, so a restore can no longer overwrite a newer write or bring back a
deleted object.

A zero-length or truncated local metadata file is now replaced from S3 with the flag on: the
filesystem provider reports it as UnreadableStoredObjectException and keeps it, and the chain
overwrites it with a readable durable copy. Without one the read fails as before, so the
metadata is never treated as absent and regenerated.

Storage gains listFirstObject, a one-key listing. The S3 provider uses it for existsGroup,
existsObject and hasDurableCopy instead of listing every page under the prefix.

With the flag on, static push and metadata clients without a region or custom endpoint use
the global endpoint and no signer override, so the SDK looks up the bucket's region and signs
SigV4 for it instead of failing to find a region. The credential-chain constructor is no
longer pinned to us-west-2 with the flag on.

The filesystem listObjectPaths resolves its prefix through normalizePath, as writes do.

The S3 provider's pushFile starts its upload inside its per-key lock, so a lock timeout no
longer leaves an upload running and pushes of one key on a node are ordered.

Unreachable flag checks in the S3 and filesystem providers are removed, and the doc now
describes flag-off credential behavior, the region handling, the corrupt-cache and locking
behavior, and that both asset-blobs/ and extracted-metadata/ are shared across namespaces.
docs/testing/BINARY_S3_STORAGE.md was not reachable from docs/README.md, which fails the docs reachability check.
@swicken
swicken force-pushed the s3-stack/1a-storage-layer branch from 801072a to aa4cc7a Compare October 7, 2026 20:47
try {
Files.copy(source.toPath(), snapshot, StandardCopyOption.REPLACE_EXISTING);
final String hash = S3ContentAddressedStorage.hash(snapshot.toFile());
final String configuration = org.apache.commons.codec.digest.DigestUtils.sha256Hex(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SharedExtractedMetadata.java:35 null extractorVersion collapses distinct parser bundles into one shared cache key

Current code:

final String configuration = org.apache.commons.codec.digest.DigestUtils.sha256Hex(
        "tika:" + extractorVersion + ":schema:" + schemaVersion + ":text-limit:" + textLimit);

Problem: TikaUtils.extractorVersion() (TikaUtils.java:672) returns null when OSGi/Tika is not initialized, so the key becomes "tika:null:..." and unrelated parser bundles share one cached extraction, violating the class's own "unknown implementations must not share extraction caches" contract.

Fix:

if (extractorVersion == null) {
    return extract.apply(source);
}
final String configuration = org.apache.commons.codec.digest.DigestUtils.sha256Hex(
        "tika:" + extractorVersion + ":schema:" + schemaVersion + ":text-limit:" + textLimit);

if (assetStorageFlag) {
// ponytail: only in-memory overrides re-read the latched flag; reloads and the system table stay latched.
com.dotcms.storage.AssetStorageFeature.reset();
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Config.java:800 drop stray 'ponytail:' token from comment

Current code:

// ponytail: only in-memory overrides re-read the latched flag; reloads and the system table stay latched.

Problem: The comment begins with an accidental leftover word that obscures its meaning.

Fix:

// Only in-memory overrides re-read the latched flag; reloads and the system table stay latched.

try {
Files.copy(source.toPath(), snapshot, StandardCopyOption.REPLACE_EXISTING);
final String hash = S3ContentAddressedStorage.hash(snapshot.toFile());
final String configuration = org.apache.commons.codec.digest.DigestUtils.sha256Hex(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SharedExtractedMetadata.java:34 guard null extractorVersion before building shared cache key

Current code:

final String configuration = org.apache.commons.codec.digest.DigestUtils.sha256Hex(
        "tika:" + extractorVersion + ":schema:" + schemaVersion + ":text-limit:" + textLimit);

Problem: TikaUtils.extractorVersion() (TikaUtils.java:672) returns null when OSGi/Tika is uninitialized, so the key becomes "tika:null:..." and unrelated parser bundles share one cached extraction, violating the class's own "unknown implementations must not share extraction caches" contract. Dead until PR 2 wires it.

Fix:

if (extractorVersion == null || extractorVersion.isBlank()) return extract.apply(source);
final String configuration = org.apache.commons.codec.digest.DigestUtils.sha256Hex(
        "tika:" + extractorVersion + ":schema:" + schemaVersion + ":text-limit:" + textLimit);

@fabrizzio-dotCMS fabrizzio-dotCMS left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the full production diff. The design matches the stack doc, and every new branch is gated, so flag-off behavior matches main.

Two things I'd like addressed before merge:

  • A local write failing after S3 accepted leaves a stale local copy that keeps being served and is never evicted (comment on ChainableStoragePersistenceAPI.pushFile).
  • The shared IdentifierStripedLock held across S3 uploads in AmazonS3StoragePersistenceAPIImpl.pushFile: harmless here, but once bundles go through it in #37774 it can fail unrelated content saves.

The rest are a verification-cost improvement in S3ContentAddressedStorage, CI coverage against a real S3 (MinIO through the existing docker-maven-plugin setup), lock hardening suggestions, a few documentation requests (target deployment, namespace, environment cloning), one question on key case, and minor items.

* node half-enabled. Set it through the environment or properties file and restart.</p>
*/
public final class AssetStorageFeature {
public static final String FLAG = "FEATURE_FLAG_S3_ASSET_STORAGE";

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FEATURE_FLAG_S3_ASSET_STORAGE should be declared in com.dotcms.featureflag.FeatureFlagName like the other flags (with a Javadoc noting it is read once per process and needs a restart), and AssetStorageFeature.FLAG should reference it, the same way IndexConfigHelper.FLAG_KEY references FeatureFlagName.FEATURE_FLAG_OPEN_SEARCH_PHASE. That keeps the flag discoverable where people look for flags.

public static final String FLAG = FeatureFlagName.FEATURE_FLAG_S3_ASSET_STORAGE;

@@ -0,0 +1,129 @@
# S3 asset storage

S3 asset storage lets dotCMS keep binary assets and their metadata durably in S3, with the

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could this page state the deployment this feature targets? Reading the parent issue, the goal is to drop the shared NFS volume: S3 holds the only durable copy and is what nodes share, and each node keeps a bounded, disposable local cache (#37868, "scale out without a shared NFS volume"). The page doesn't say that, and a reader can easily conclude the opposite: that the flag adds S3 on top of an existing NFS asset directory, which only adds a layer (and eviction on a shared NFS directory is called out as unsafe later in the stack).

Two things would help:

  • Say explicitly that the intended setup is a node-local asset directory per node, and whether running with the flag on over a shared NFS directory is supported, discouraged or untested.
  • Say whether serving directly from S3 (streaming, or range requests without a full download to the local cache) is planned, or out of scope. Today a cold read downloads the whole object before serving, including for range requests on large files.

private final List<StoragePersistenceAPI> storagePersistenceAPIList;
private final ObjectWriterDelegate defaultWriterDelegate;
private final Chainable404StorageCache cache;
private final com.google.common.util.concurrent.Striped<java.util.concurrent.locks.Lock> assetLocks =

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Question on key case. The lock key is groupName + "/" + path, case-sensitive, and S3 keys keep their case (transformReadPath never lowercases with the flag on). The filesystem provider, though, lowercases every path (FileSystemStoragePersistenceAPIImpl.normalizePath). So two keys that differ only in case (.../File.json and .../file.json) are two S3 objects and take two different locks, but share one local file. The restore/delete ordering this lock guarantees doesn't hold for that pair, and the local cache can return one key's bytes for the other.

In this PR that only reaches the metadata chain, where keys come from inodes and field names, so it may be impossible in practice. If so, could that be stated here? Otherwise, deriving the lock key with the same normalization the local provider applies would close it. (1b keeps case locally for binary-assets/generated-assets, but dotmetadata stays lowercased.)

private final ObjectWriterDelegate defaultWriterDelegate;
private final Chainable404StorageCache cache;
private final com.google.common.util.concurrent.Striped<java.util.concurrent.locks.Lock> assetLocks =
com.google.common.util.concurrent.Striped.lock(256);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: the package names are spelled inline here instead of imported.


if (AssetStorageFeature.isEnabled()) {
final var lock = assetLocks.get(groupName + "/" + path);
lock.lock();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two suggestions on the per-key lock, same pattern in all five flag-on branches (145, 254, 288, 375, 472):

  1. Make assetLocks static. The lock protects a key, not an instance, but each chain instance has its own stripes. Today each group is owned by one effectively-singleton chain (metadata via StoragePersistenceProvider, binaries via BinaryAssetStorageAPIImpl in 1b, bundles in 4), so it works by discipline. StoragePersistenceProvider.forceInitialize() rebuilds the chain, after which an in-flight operation on the old instance and a new one on the new instance no longer exclude each other. A process-wide field removes that dependency at almost no cost (raise the stripe count if needed). Worth a comment that code holding one of these locks must never call into another chain, since there is no timeout.

  2. Bound the wait. lock() waits forever and is not interruptible, and the critical section does real I/O: S3 transfers (the SDK only cuts after 50 s without data; a slow transfer that keeps progressing has no cap) and local writes, which on a hard-mounted NFS directory can block indefinitely. One stuck holder then parks every thread that needs a key on that stripe, which can exhaust request threads. tryLock with a generous timeout (minutes) that fails the operation with a clear error would keep the correctness guarantee (it never proceeds without the lock) while making a stuck holder visible.

if (!lock.tryLock(LOCK_WAIT_MINUTES, TimeUnit.MINUTES)) {
    throw new DotDataException("Timed out waiting for storage key " + groupName + "/" + path);
}

try {
storage.uploadFileIfAbsent(bucketName, transformReadPath(groupName, path), file);
} catch (com.amazonaws.services.s3.model.AmazonS3Exception conflict) {
if (conflict.getStatusCode() != 412) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: here (and in backfillObject, line 856) only 412 counts as "someone else wrote first", while S3ContentAddressedStorage.preconditionFailure and AWSS3Storage.uploadFileIfMatch also accept 409. AWS returns 409 ConditionalRequestConflict when two conditional writes to the same key race. With a 409 the backfill throws instead of falling through to the hasDurableCopy re-verification below, so the batch fails and relies on the job retry. Treating 409 like 412 here (or sharing one preconditionFailure helper) would make the three paths consistent.

}
final String key = transformReadPath(groupName, path);
try (final InputStream input = Files.newInputStream(file.toPath())) {
final String md5 = DigestUtils.md5Hex(input);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: the MD5 reads the whole local file before we know S3 has the key or that the size matches; when either fails the answer is false without any hash. For backfill of large files not yet in S3 that is a full read for nothing. Moving the digest after the listFirstObject and size check doesn't change the result.

final var object = storage.listFirstObject(bucketName, key);
if (object == null || !key.equals(object.getKey()) || object.getSize() != file.length()) {
    return false;
}
try (InputStream input = Files.newInputStream(file.toPath())) {
    return DigestUtils.md5Hex(input).equalsIgnoreCase(object.getETag())
            || storage.fileContentsMatch(bucketName, key, file);
}

}
consumeReference(published);
}
if (!matches(ownerKey, snapshot.toFile())) throw new DotDataException("Shared asset bytes were not verified after publication");

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Each store downloads the full blob twice to verify it: blobMatches after the conditional upload (line 84) and matches here, which also re-hashes the local snapshot. A new blob costs one upload plus two full downloads; a deduplicated one costs two full downloads. From #37772 this runs inside the check-in transaction (storeRevision under ESContentletAPIImpl.checkin), with the contentlet row locked, so a 1 GB binary keeps the transaction open for the upload plus 2 GB of reads.

Only one of those comparisons protects against something:

  • After an upload that succeeded, the SDK has already validated the transferred bytes, so re-reading the blob adds nothing. The byte comparison matters when the blob already existed (412/409 on uploadFileIfAbsent, or blobMatches true on entry), because then we don't know who wrote it.
  • matches at the end repeats that same blob comparison. What it adds is the reference header check, which the published block just above already does.

Suggestion: compare blob bytes only when the blob pre-existed, and end with the reference check only. A new blob then needs no download and a deduplicated one needs one, with the same guarantees.

import static org.mockito.Mockito.*;

/** Real filesystem + AWS adapter + S3 server. See docs/testing/BINARY_S3_STORAGE.md. */
@EnabledIfSystemProperty(named = "s3.test.endpoint", matches = ".+")

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This class is the only test in the PR that talks to a real S3 API, and CI never sets s3.test.endpoint, so it is always skipped there. The rest of the suite (notably AssetStorageFeatureTest) runs in CI with mocks and fakes, which covers the chain and provider logic well, but the behavior that depends on a real server never runs in CI: If-None-Match/If-Match handling, listObjects pagination, and the real 404 NoSuchKey/409/412 responses that pullFile, uploadFileIfAbsent and uploadFileIfMatch map. Per the PR description, this is also the only PR in the stack that gets the full CI run.

Suggestion: add MinIO as one more entry in the docker-maven-plugin imagesMap in parent/pom.xml, next to database and opensearch, and pass -Ds3.test.endpoint plus the bucket properties to the test runs. That follows the pattern already used for Postgres and OpenSearch, works the same locally and in CI, and would let this class and the flag-on integration tests (-Ds3.cms.enabled=true) run on every PR of the stack. Untested paths in this PR that would benefit: S3ContentAddressedStorage, the multipart branch of uploadFileIfAbsent (above 32 MiB), and a local write failing after S3 accepted.

return value;
}

private String groupKey(final String group) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Docs suggestion on the namespace. The prefix is added here, while what the database stores (e.g. storageKey from #37772) is the key without it, so the configured value has to stay the same for an installation's lifetime and be identical on every node that shares its database. If a node reads a different value (or none), its writes and reads land under another prefix and content looks missing on that node.

dotmarketing-config.properties already warns that changing it "requires migration and separate local cache roots". It would help to say in BINARY_S3_STORAGE.md:

  • every node of one installation (one database) must use the same value;
  • like the flag, it shouldn't be set only in the system table: Config gives an environment variable precedence over the system table, and if this provider is built before the system table source is initialized, it falls back to the properties file (empty) and keeps that value until restart, since namespace is final;
  • how to clone an environment. Copying production's database into staging doesn't work on its own: with its own namespace (needed so staging's cleanup jobs don't delete production's objects), staging's rows point at keys that only exist under production's prefix, so it sees no binaries; with the same namespace, staging's deletions remove production's binaries. The supported path is starter export/import (feat(storage): add S3 binary backfill, starter export/import and integrity repair #37773), which publishes the assets under the target installation's namespace before the import commits.

Optional, if you think it's worth it: record the namespace in use on first enable (e.g. a database row) and refuse to start when the configured value differs.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

dotbot code review:

  • Reviewer: meta/muse-spark-1.3 (medium)
  • Overall: patch is incorrect
  • New findings this run: 0
  • Prior unresolved dotbot findings still relevant: 1
  • Active findings total: 1

No new actionable bugs were found in the current changes, but 1 prior unresolved dotbot finding still applies, so the patch remains incorrect.

Tip: comment with "/dotbot address comments" to attempt automated fixes for unresolved review threads.

reviewed by dotbot · meta/muse-spark-1.3 · medium

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

spec.md:456 resolve FR-012 contradiction between 260px minimum and shrink-to-fit

Posted as a general PR comment because the referenced file is not part of this PR's diff.
Original target: specs/37930-content-drive-grid-view/spec.md:454-457.

Current code:

- **FR-012**: Cards MUST be at least 260px wide, as in Content Search's card view. The grid MUST fit
  as many columns as that allows, ... and never scroll horizontally. Below 260px of
  available width it MUST show one card that shrinks to fit.

Problem: A card that "shrinks to fit" below 260px contradicts "MUST be at least 260px wide" in the same requirement; an implementer cannot satisfy both.

Fix:

- **FR-012**: Cards MUST be at least 260px wide when the available width allows. ... Below 260px of
  available width it MUST show one card that shrinks to fit.

@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

dotbot code review:

  • Reviewer: ~z-ai/glm-latest (medium)
  • Overall: patch is incorrect
  • New findings this run: 1
  • Prior unresolved dotbot findings still relevant: 1
  • Active findings total: 2
  • Findings remapped to general PR comments: 1 (missing file map=1)

1 new actionable finding was identified in the current changes, and 1 prior unresolved dotbot finding still applies, so the patch remains incorrect.

The incremental delta adds only the Content Drive grid-view specification; no production code changed since the previously reviewed head. The spec is consistent with the existing keybindings spec (#32591) except for a minor wording contradiction in FR-012, a P3 spec-edit item. The prior Config.java comment typo remains unfixed in the branch and is carried forward.

Tip: comment with "/dotbot address comments" to attempt automated fixes for unresolved review threads.

reviewed by dotbot · ~z-ai/glm-latest · medium

@claude

claude Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Pull Request Unsafe to Rollback!!!

  • Category: H-5 — Binary Storage Provider Change
  • Risk Level: 🟠 HIGH (conditional: only manifests if the new opt-in flag is turned on and used before a rollback)
  • Why it's unsafe: This PR adds an opt-in S3 asset lifecycle gated by AssetStorageFeature (new file dotCMS/src/main/java/com/dotcms/storage/AssetStorageFeature.java), wired to FEATURE_FLAG_S3_ASSET_STORAGE=false in dotCMS/src/main/resources/dotmarketing-config.properties (default off, read once at startup). When the flag is turned on, AmazonS3StoragePersistenceAPIImpl.transformReadPath()/groupKey() switch binary assets to a new key layout (optionally namespaced via storage.file-metadata.s3.namespace), and new class S3ContentAddressedStorage (dotCMS/src/main/java/com/dotcms/storage/S3ContentAddressedStorage.java) writes a small JSON pointer ({\"sha256\":...,\"length\":...}) at the asset's owner key instead of raw bytes, with the real bytes stored separately under asset-blobs/sha256/<hash prefix>/.../<hash>. N-1 has no AssetStorageFeature class and always uses the legacy path layout — after a rollback it will either look for assets at paths that no longer match what N wrote (404s) or, if it does hit the owner key, hand back the small JSON reference string as if it were the real binary (corrupted downloads/renditions) for any asset written while the flag was enabled.
  • Code that makes it unsafe:
    • dotCMS/src/main/java/com/dotcms/storage/AssetStorageFeature.java (new feature flag, default false)
    • dotCMS/src/main/resources/dotmarketing-config.properties (FEATURE_FLAG_S3_ASSET_STORAGE, storage.file-metadata.s3.namespace)
    • dotCMS/src/main/java/com/dotcms/storage/AmazonS3StoragePersistenceAPIImpl.java (transformReadPath, groupKey, pushFile, pullFile branch on AssetStorageFeature.isEnabled() to use the new layout)
    • dotCMS/src/main/java/com/dotcms/storage/S3ContentAddressedStorage.java (new content-addressed blob/reference format, store()/retrieve())
    • dotCMS/src/main/java/com/dotcms/storage/ChainableStoragePersistenceAPI.java (routes pushFile/pullFile through the durable S3 chain when enabled)
  • Alternative (if possible): This is already partially mitigated by the two-phase pattern (feature defaults off, and S3ContentAddressedStorage.retrieve() has a built-in fallback that reads pre-existing raw objects when no SHA-256 marker is present, which protects forward migration). To make it rollback-safe as well: document in the release notes that enabling FEATURE_FLAG_S3_ASSET_STORAGE is a one-way operational decision — once assets have been written under the new content-addressed/namespaced layout, do not roll the binary back below this release without first backfilling those assets back to the legacy flat S3 layout.

@jcastro-dotcms jcastro-dotcms left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the detailed write-up and the docs, they made this a lot easier to review. I've been doing the QA and code review of this PR focused on one question: does dotCMS behave exactly like main with the flag off?

Short answer: yes, with one documented exception. I traced every changed method in the chain, filesystem, database and S3 providers, FileStorageAPIImpl, and the static push publishing classes, and with the flag off each one falls back to main's code. I also ran the tests locally on the PR head:

  • the PR's own unit tests: 24 passed, 2 skipped (the MinIO ones, see the doc comment);
  • PublishingEndPointFactoryTest: 5 passed;
  • the existing integration tests that cover the changed code, with the flag off: FileStorageAPITest, StoragePersistenceAPITest and FileMetadataAPITest all pass. PublishingEndPointTest has 2 failures, but they fail the same way on main when the class runs on its own (the endpoint's catch block calls PortalUtil.getUser().getUserId() with no request on the thread). That's a pre-existing test-isolation issue unrelated to this PR, and it passes in CI inside MainSuite1a.

I agree with the points Fabrizzio raised, especially the stale local copy after S3 accepts a write, and the shared IdentifierStripedLock held during uploads. The inline comments below cover only what his review doesn't already mention:

  1. The STS module is the only flag-off behavior change, and I'd like it documented and in the release notes (dotCMS/pom.xml).
  2. With the flag on and S3_STORAGE_FILE_REPO_TYPE set to LOCAL/HASH_LOCAL, concurrent downloads of the same key collide and downloaded files are never cleaned up (AmazonS3StoragePersistenceAPIImpl).
  3. The MinIO image in the doc can't be pulled, so BinaryS3StorageTest can't run anywhere (BINARY_S3_STORAGE.md).
  4. Two smaller ones in Config: a leftover ponytail: in a comment, and Config now depending on the storage feature.

I'm requesting changes for 1–3 together with Fabrizzio's two blocking items. 4 is minor.

The pre-merge manual test plan for the flag-off scenarios will follow as a separate comment on this PR.

Logger.info(Config.class, "Setting property: " + key + " to " + value);
props.setProperty(key, value);
if (assetStorageFlag) {
// ponytail: only in-memory overrides re-read the latched flag; reloads and the system table stay latched.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Small one: this comment starts with the word ponytail:, which doesn't mean anything in this context and looks like text left over from an earlier draft or a note-to-self. The rest of the sentence is accurate and useful, so it's only the leading token that should go.

It's harmless at runtime, but Config is one of the most-read classes in the codebase, and a reader will stop and wonder whether ponytail refers to some mechanism they should know about. Here's the same comment without it:

Suggested change
// ponytail: only in-memory overrides re-read the latched flag; reloads and the system table stay latched.
// Only in-memory overrides re-read the latched flag; reloads and the system table stay latched.

*/
public static void setProperty(String key, Object value) {
if (props != null) {
final boolean assetStorageFlag = com.dotcms.storage.AssetStorageFeature.FLAG.equals(key);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is about which class depends on which, not about whether it works (it does).

Config is one of the lowest-level utilities we have: it lives in com.dotmarketing.util and pretty much everything else depends on it. With this change, setProperty now checks for one specific feature's key and calls com.dotcms.storage.AssetStorageFeature.reset() directly, written fully qualified in two places. So the generic configuration class now knows about the S3 storage feature, which is the opposite of the usual direction.

Why it matters beyond style:

  • If another flag later needs the same "read once, reset in tests" behavior, the natural move is to add another if here, and Config slowly collects feature-specific special cases.
  • Someone reading Config.setProperty has no hint of why storage code appears there unless they already know the latching design.

A couple of options, in order of preference:

  1. Give Config a small, generic hook for in-memory overrides (for example, a way to register a callback for a given key), and have AssetStorageFeature register its own reset. Config then stays feature-agnostic.
  2. If that's more than you want for this PR, keep it as is but import the class and add a one-line comment explaining that the S3 flag is latched at first read, and that this reset only exists so tests can switch modes.

Not a blocker on its own, but I'd like to see one of the two.

Comment thread dotCMS/pom.xml
</dependency>
<dependency>
<groupId>com.amazonaws</groupId>
<artifactId>aws-java-sdk-sts</artifactId>

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I want to flag this one specifically, because it's the only change in the PR that reaches installations running with the flag off, and the goal of this review is that flag-off behaves exactly like main.

What I verified. Adding the STS module changes how the AWS SDK v1 resolves credentials when no access key and secret are configured. NoWebIdentityCredentialsProviderChain handles this well for dotCMS's own code: the three places where dotCMS falls back to the default chain (the S3 metadata provider, static push publishing, and the static push endpoint validation) use a chain without the web-identity step while the flag is off. Those three keep resolving credentials exactly as they did on main.

What still changes with the flag off (the PR description mentions both, but only briefly):

  1. AWS profiles that use role_arn. On main, when the SDK reaches a profile with role_arn, it tries to load STS by reflection, fails with "To use assume role profiles the aws-java-sdk-sts module must be on the class path" (I confirmed the message in the SDK jar), and quietly moves on to the next provider, usually the EC2/ECS instance role. With STS packaged, the role is now actually assumed. So the identity dotCMS uses for S3 can change after an upgrade, with no configuration change on the customer's side. If the assumed role and the instance role have different S3 permissions, static push publishing or the S3 metadata provider could start failing with AccessDenied, or start succeeding where they used to fail.
  2. Plugins that build their own DefaultAWSCredentialsProviderChain with the SDK classes dotCMS provides. In an EKS pod with a web identity token, those plugins now get the pod's role instead of the node's role.

What it doesn't affect: installations with static keys configured (static keys always win before these steps), SDK v2 clients such as the SQS code (v2 has its own separate STS module), and stored data. Rolling back just removes the jar and returns to the old resolution, so there's no rollback-safety concern.

What I'm asking for: the risk is low, and both setups are uncommon in our Docker/Kubernetes deployments. But it's the one flag-off difference, and if it ever bites someone it will look like an unrelated permissions problem right after an upgrade. Could you:

  • add a short note in BINARY_S3_STORAGE.md under "Behavior with the flag off", and
  • make sure it goes into the release notes for the release that ships this PR?

That way support has something to point to if a customer reports S3 AccessDenied after upgrading.

@EnterpriseFeature(licenseLevel = LicenseLevel.PLATFORM, errorMsg = INVALID_LICENSE)
public File pullFile(final String groupName, final String path) throws DotDataException {
if (AssetStorageFeature.isEnabled()) {
final File download = fileRepositoryManager.getOrCreateFile(path);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a flag-on issue that only appears with a non-default setting, but it's easy to miss.

With the flag on, pullFile downloads the object into fileRepositoryManager.getOrCreateFile(path). Which file that is depends on S3_STORAGE_FILE_REPO_TYPE:

  • With the default (TEMP), each call gets its own temporary file, and releaseRetrievedFile (line 641) deletes it afterwards. That works correctly.
  • With LOCAL or HASH_LOCAL, the file is a fixed path derived from the key. Two problems follow from that.

1. Concurrent downloads of the same key collide. If two threads on the same node call pullFile for the same key at the same time, both download into the same file at once and can produce a mixed or truncated result. The chain's assetLocks prevents this when the call comes through ChainableStoragePersistenceAPI, but this method has no lock of its own. So anything that calls the S3 provider directly (the remoteObjectStorage() users later in the stack), or two separate chain instances, aren't protected.

2. Downloaded files are never cleaned up. releaseRetrievedFile only deletes when the repo is a TempFileRepositoryManager. With LOCAL/HASH_LOCAL, after the chain restores an object into the filesystem provider, the downloaded file stays where it is. Every restored asset ends up stored twice on the node's disk: once in the asset directory and once in the repo directory. For a feature whose point is a bounded local cache, that works against the goal, and eviction (in 1b) won't know about the second copy.

Suggestion: with the flag on, always download to a unique temporary file (Files.createTempFile, the same approach pushObject already takes when the flag is on), whatever the repo type, and always delete it in releaseRetrievedFile. If LOCAL/HASH_LOCAL aren't meant to be supported with the flag on, the alternative is to reject them at startup with a clear message and say so in BINARY_S3_STORAGE.md, so nobody runs into this in production.

-p 127.0.0.1:19002:9000 \
-e MINIO_ROOT_USER=binary-storage-test \
-e MINIO_ROOT_PASSWORD=binary-storage-test \
minio/minio:latest server /data

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried to follow this section to run the MinIO tests locally during QA, and couldn't, because the image can't be pulled:

  • docker pull minio/minio:latest fails with pull access denied for minio/minio, repository does not exist or may require 'docker login'
  • docker pull quay.io/minio/minio:latest (MinIO's own registry) fails with 401 UNAUTHORIZED

Other Docker Hub images pull fine on the same machine, so it's this image specifically, not a general Docker or network problem.

Why it matters: BinaryS3StorageTest is the only test in this PR that talks to a real S3 API, and CI already skips it because there's no endpoint there. If developers can't run it locally either, nobody exercises it: the conditional writes (If-None-Match / If-Match), the real 404/409/412 handling, pagination, and S3ContentAddressedStorage against a real server. As a result, my local run of the PR's unit tests had to skip both tests in that class.

What I'm asking for:

  • Point the doc at an image that can actually be pulled, ideally pinned to a specific tag or digest so it doesn't silently change. A dotCMS-mirrored image would be ideal.
  • Use that same image for the docker-maven-plugin entry Fabrizzio suggested in his BinaryS3StorageTest comment, so the tests run identically in CI and locally.
  • If MinIO is replaced with another S3-compatible server, it needs to support conditional PUTs (If-None-Match: * and If-Match), because these tests depend on them.

Once there's a working image I'll re-run the 2 skipped tests and report back on this PR.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

AI: Not Safe To Rollback Area : Backend PR changes Java/Maven backend code Area : Documentation PR changes documentation files PR : dotbot review Trigger dotbot AI code review and the post-merge QA test plan

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

4 participants