Repository navigation
Replies: 1 comment
|
I've added Change Data Capture mixed index synchronization into JanusGraph version 1.2.0 (to be released), but this is specifically targeting Cassandra storage backend and it wasn't added to any other storage backends yet. CDC Mixed Index Synchronization landed in #4906 . |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
What I'm asking
Whether upstream has appetite for a durable, deferred mixed-index write mode, and if so
what shape it should take. I'd rather find out before writing code, because the natural
implementation touches the storage capability contract.
I'm not proposing a specific API yet. I mostly want to check that the problem framing below
matches how maintainers see it, and hear if anyone has already gone down this road.
The current behaviour
StandardJanusGraph.commit()commits primary storage, then writes mixed indexes from the samethread, and collects rather than propagates any failure. The code says so directly:
I don't think this is a bug — it looks like a deliberate choice, since by that point primary
storage is durably committed and cannot be rolled back. But it leaves the write in a position
that is both too late to abort and too early to hand off:
IndexTransaction.mutationsis nulled influshInternal, andreleaseTransaction()runs ina
finally. Oncecommit()returns, the delta needed to retry no longer exists in theprocess.
StandardTransactionLogProcessor.restoreExternalIndexes— but requirestx.log-tx=true(defaultfalse) and an operator-started recovery process. Out of thebox, a failed index write is permanent, and the only signal is one ERROR line.
Three consequences follow, and I'd like to know whether others see them the same way.
1. Transient backend errors are treated as permanent. For Elasticsearch,
convert()mapseverything except
InterruptedExceptiontoPermanentBackendException, andBackendOperationonly retriesTemporaryBackendException— so a 429 or a socket timeout isdropped without using the retry loop that is already wired up. I filed that separately as
#4925, along with #4926 for a related case where 404 bulk items are discarded. Those two are
ordinary bugs and fixable in place; they reduce the rate but not the class.
2. A process death between the two commits loses the write with no trace. The pending
index mutation exists only in heap between
commitStorage()andcommitIndexes(). OnKubernetes this window is hit routinely during rollouts. There is no exception, so nothing
logs and no metric moves.
3. The index write is on the request thread, and its retry budget is per index backend.
BackendOperation.executeDirectinvokes the callable directly and sleeps in place for backoff,up to
storage.write-time(default 100 s). A degraded index backend therefore parks requestthreads, which on a bounded pool becomes thread-pool exhaustion — so an index-backend brownout
can stop the graph serving reads.
Worth being explicit that this is specific to mixed indexes. Composite index entries go to
the index store and vertex-centric index columns to the edge store, both inside the same
StoreTransaction, so they commit atomically with the data and none of this applies to them.Why retrying harder doesn't close it
Retry helps case 1 and does nothing for case 2, and case 3 gets worse with a longer budget.
Making the client retry doesn't work either: once storage has committed, a client retry is
either a no-op or a duplicate depending on whether the caller's mutation is idempotent, and it
cannot repair the index in the common case where the graph already reflects the desired state.
What seems to be needed is somewhere durable to put the outstanding index work at the moment
the storage commit succeeds.
The shape I keep arriving at
Sketching this to make the question concrete, not to propose it as-is.
Two properties seem important:
Carry document identities, not materialised documents or field deltas. If the record holds
a document, a replay after a schema change writes the field set that existed at commit time
rather than now; and index applicability itself isn't stable, since
indexOnlyconstraintsdepend on the element's label. Resolving both at processing time — from the element plus the
live schema — makes delayed and duplicate delivery safe. It also rules out versioned deltas:
if v42 reaches the backend before v41, rejecting v41 is only correct when v42 is complete.
Reuse the existing pieces rather than adding parallel ones. As far as I can tell most of
what's needed is already there:
IndexTransactionalready groups mixed-index mutations by backing index, store, anddocument id.
IndexSerializer.reindexElementalready builds a current document for an element and index.IndexProvider.restorealready replaces a complete document, and deletes when handed anempty one — which incidentally fixes the case where removing an element's last indexed
property leaves a stale document behind.
BackendTransaction.getTxLogPersistor()/ExternalCachePersistoralready allow a log entryto join the same
CacheTransaction, andcommit()already stagesPRIMARY_SUCCESSthroughthe primary transaction before
commitStorage().So the genuinely new parts look like: a self-contained committed index-work record, a
direct/deferredper-index-backend write mode withdirectas the default, and a way for abackend to declare that it can commit graph mutations and that record atomically.
The awkward part, and the actual question
StoreFeatures.hasTxIsolation()doesn't say anything about whether mutations across the graphstores and a log store share a commit outcome. A
deferredmode is only sound on a backendthat guarantees they do, so it seems to need a new capability — something like
supportsAtomicMultiStoreCommit(), defaulting tofalseso no existing backend silentlyacquires a stronger contract, with configuration validation refusing
deferredrather thandowngrading silently.
That's the part I don't want to guess at:
deferredwrite mode something upstream wants at all, or is the position thatindex durability belongs entirely to the deployer via
tx.log-txplus a recovery process?preferred way to express "these stores commit together" that I've missed?
separate logical store? The log's write side looks reusable; its read side seems less so —
KCVSLog.prepareMessageProcessingsubmits callbacks fire-and-forget,ProcessMessageJobcatchesThrowableand only logs, andsetReadMarkeradvances thecheckpoint regardless of whether any callback succeeded. So an always-on consumer would need
acknowledgement-aware reading, which is a second design question.
Happy to write it up properly as a design proposal, and to do the implementation in reviewable
steps, if the direction is one maintainers would accept.
All reactions