Skip to content

[SPIKE] Migrate Experiments from legacy (CubeJS) to new analytics infrastructure (CAEM + ClickHouse) #36195

Description

@erickgonzalez

Research Question

What is the effort and ETA to migrate the Experiments feature off the legacy analytics infrastructure (CubeJS-based) and onto the new analytics infrastructure (CAEM + ClickHouse)?

The spike must produce a concrete migration plan that accounts for every affected layer and answers the open questions below.

Scope to investigate

  • Core — Experiments backend logic, services, and current dependencies on the legacy analytics/CubeJS pipeline.
  • Apps — any app-level integrations or configuration tied to Experiments data collection/reporting.
  • Frontend — Experiments UI (configuration, results/reporting views) and its current data source(s).
  • CAEM — what the new infrastructure provides today and what gaps exist for Experiments use cases.
  • ClickHouse — schema/tables needed for Experiments event and results data on the new pipeline.
  • Endpoints — design and creation of the new REST endpoints that replace CubeJS queries (Experiments is no longer using CubeJS).

Key open questions

  1. CubeJS replacement — which queries/aggregations does Experiments rely on today via CubeJS, and how are they reproduced as new endpoints over ClickHouse/CAEM?
  2. Data migration — Experiments is in use by a few customers. Is it feasible to migrate existing experiment data from the old pipeline to the new one? At what cost/risk? (Evaluate feasibility — not a hard requirement.)
  3. BE data access for customers — can customers inspect their Experiments data directly from the backend on the new infrastructure? Determine whether this is possible or not (exploratory, not a requirement).

Timebox

8h

Acceptance Criteria

  • Inventory of legacy dependencies — document every Experiments touchpoint on the old infrastructure (CubeJS queries, data sources, schemas) across core, apps, and frontend.
  • Target architecture — describe the new flow on CAEM + ClickHouse, including the ClickHouse schema/tables required for Experiments.
  • Endpoint design — list the new REST endpoints that must be created to replace CubeJS (purpose, inputs, outputs, ownership), with a rough complexity estimate per endpoint.
  • Per-layer effort breakdown — estimate work for core, apps, frontend, CAEM, ClickHouse, and endpoints, with an aggregated ETA and effort.
  • Data migration assessment — document whether migrating existing customer experiment data from old → new is feasible, with approach options, risks, and a recommendation (or a clear "not feasible / not worth it" with reasoning).
  • Customer BE data access feasibility — answer whether customers can check their Experiments data from the BE on the new infrastructure: possible or not, with constraints.
  • Risks & sequencing — identify migration risks (customer impact, data continuity, dual-running period) and a recommended phasing/rollout order.
  • Recommendation — a prioritized migration plan with the final ETA/effort the team can use for planning.

Context

Experiments currently runs on the legacy analytics infrastructure, querying data through CubeJS. The platform is standardizing on the new analytics infrastructure (CAEM + ClickHouse), and CubeJS is being retired. Before committing to the migration, the Falcon team needs a clear picture of the effort, the new endpoints required, and the implications for the few customers already using Experiments — particularly around data continuity (migrating historical experiment data) and the ability to access experiment data directly from the backend.

This is a timeboxed spike to turn these unknowns into a concrete, estimated plan.

Metadata

Metadata

Assignees

Type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions