Skip to content

fix: dispatch map lookups with normalized keys and nondeterministic null-guarded children #6056

fix: dispatch map lookups with normalized keys and nondeterministic null-guarded children

fix: dispatch map lookups with normalized keys and nondeterministic null-guarded children #6056

Workflow file for this run

# Licensed to the Apache Software Foundation (ASF) under one
# or more contributor license agreements. See the NOTICE file
# distributed with this work for additional information
# regarding copyright ownership. The ASF licenses this file
# to you under the Apache License, Version 2.0 (the
# "License"); you may not use this file except in compliance
# with the License. You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing,
# software distributed under the License is distributed on an
# "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
# KIND, either express or implied. See the License for the
# specific language governing permissions and limitations
# under the License.
# Top-level CI orchestrator: runs cheap preflight checks first, then fans out
# to the long-running test/build workflows only if preflight passed and the
# change touched files relevant to that workflow.
#
# Merging goes through GitHub's merge queue (see `rulesets` in `.asf.yaml`), so
# there are three tiers:
#
# pull_request fast feedback. The Linux build only, with its test matrix
# run against the default Spark profile (4.1) alone.
# merge_group the authoritative gate. The PR tier plus the macOS build,
# the benchmark compile check, the Delta contrib build gate,
# the PyArrow UDF suite, Spark SQL on Spark 4.1 and Iceberg
# 1.11, evaluated against the merge result rather than the
# PR head. One Spark version and one Iceberg version, both
# the default profile's.
# schedule the nightly regression sweep of everything else: the Linux
# test matrix against the other Spark profiles, Spark SQL on
# Spark 3.5 and 4.0, and Iceberg 1.8/1.9/1.10, routed by the
# same path filters over what landed on main since the last
# green nightly. A failure opens (or comments on) an issue
# labelled `ci-nightly-failure`.
#
# Spark 3.4 is deprecated and sits outside all three tiers: it runs only when
# a pull request carries `run-spark-3.4-tests`, or from a manual dispatch.
#
# Which tier a job sits in is POLICY in dev/ci/compute-changes.py, not an
# expression here. Heavy jobs deliberately have no `push` tier: the queue
# already tested the exact commit that lands, so re-running the pipeline on
# push to main would double the cost of every merge. Two exceptions: `docs`
# deploys to asf-site and must run after the commit is on main, and the Linux
# build runs on push so that main's actions/cache entries stay fresh (caches
# saved on the queue's throwaway branch are deleted with it). That second
# exception runs in `cache-refresh-only` mode: the cache writers and nothing
# else, because the queue has already tested the tree that landed.
# Project-qualified on purpose. Several Apache DataFusion repositories -- the
# core project and the other subprojects -- have a top-level workflow, and an
# unqualified `CI` is indistinguishable from theirs anywhere runs from more
# than one repository are listed together: the personal Actions dashboard,
# notification emails, and cross-repository search. Nothing depends on this
# string; every reference to this workflow, in `.asf.yaml`, `dev/ci/`, and the
# docs, goes through the file name, and the required status check is the
# `Required Checks` job name below.
name: Comet CI
# A `labeled` event (e.g. the run-spark-*-tests gates, or dependabot's automatic
# `dependencies` label added ~1s after open) fires at the same commit as the
# opened/synchronize run. Label runs reach this file through ci_label.yml, so
# `github.workflow` already separates them from the commit run; keying the
# group on the label name as well keeps two different labels from cancelling
# each other. Every other run is keyed on its event name: opened, synchronize
# and reopened all map to `pull_request`, so a new push still supersedes its
# predecessor, while a scheduled run at the tip of main never cancels that
# commit's push run.
concurrency:
group: ${{ github.repository }}-${{ github.head_ref || github.sha }}-${{ github.workflow }}-${{ github.event.action == 'labeled' && github.event.label.name || github.event_name }}
cancel-in-progress: true
# No `labeled` here: ci_label.yml handles that event and calls this file
# through `workflow_call`. When one workflow runs twice at the same commit,
# GitHub evaluates the pull request's required checks against only one of the
# two runs. A PR opened with a label already applied gets both an `opened` and
# a `labeled` run, and when the label run was the one GitHub picked, the
# `Required Checks` context showed as "Expected" forever and the merge queue
# never took the PR (issue #6159). A separate workflow gets its own check
# suite, so a label run can never hide the commit run's verdict.
# dev/ci/check-ci-config.py enforces this.
on:
pull_request:
types: [opened, synchronize, reopened]
workflow_call:
merge_group:
push:
branches:
- main
schedule:
# 06:00 UTC daily: after the evening's merges in the Americas and before
# the working day starts in Europe, so a red nightly is waiting when the
# first people look. Miri runs at 04:00 and the snapshot publish at 03:00.
- cron: '0 6 * * *'
workflow_dispatch:
jobs:
# ---------------------------------------------------------------------------
# preflight: cheap checks that gate everything else. Failure short-circuits
# the entire pipeline before any heavy job spins up. Folds in what used to be
# pr_rat_check, pr_markdown_format, pr_missing_suites, and validate_workflows.
# pr_title_check stays a standalone workflow because it needs to fire on PR
# `edited` events.
#
# This job deliberately carries no `if:`. A job held back by `if:` still
# publishes a check run under its own name with conclusion `skipped`, and the
# newest check run for a name is the one the merge box, `gh pr checks` and
# required-status-check evaluation read. Guarding this job on the label name
# therefore let any non-gating label overwrite the commit run's real
# `Preflight` verdict with `skipped` (issue #5007). Running it unconditionally
# costs about a minute of ubuntu-slim time per label event and keeps the
# reported verdict truthful. Skipping the redundant work is the heavy jobs'
# job; POLICY in dev/ci/compute-changes.py drops every job a `labeled` event
# does not gate.
# ---------------------------------------------------------------------------
preflight:
name: Preflight
runs-on: ubuntu-slim
steps:
- uses: actions/checkout@v7
- name: Set up Java
uses: actions/setup-java@v6
with:
distribution: temurin
java-version: 17
# Preflight gates every other job, so a single failed download of the
# Maven distribution here would fail the whole run.
- name: Bootstrap Maven
uses: ./.github/actions/maven-bootstrap
- name: Apache RAT license check
run: ./mvnw -B -N apache-rat:check
- name: Setup Node.js
uses: actions/setup-node@v7
with:
node-version: '24'
- name: Install prettier
run: npm install -g prettier
- name: Check markdown formatting
run: prettier --check "**/*.md"
# The site draws its ```mermaid fences with mmdc at build time, and
# sphinxcontrib-mermaid turns any render failure into a warning and a
# dropped diagram, so the deploy stays green and the page publishes with
# a hole in it (issue #6062). The docs job runs only on push to main, so
# rendering the fences here is the only chance to catch that before it
# publishes. Kept identical to the same pair of steps in docs.yaml.
- name: Install mermaid-cli
run: |
set -eux
# ubuntu-slim carries none of Chrome's shared libraries, so the browser downloads
# fine and then dies with `libatk-1.0.so.0: cannot open shared object file`. These
# are what chrome-headless-shell actually links against; libgtk-3 pulls in most of
# the rest (atk, cairo, pango, gdk-pixbuf) as dependencies.
sudo apt-get update -qq
sudo apt-get install -y -qq --no-install-recommends \
libgtk-3-0t64 libnss3 libasound2t64 libgbm1 libatk-bridge2.0-0t64 \
libcups2t64 libxkbcommon0 libxdamage1 libxcomposite1 libxrandr2 libxfixes3
npm install --prefix "$RUNNER_TEMP/mermaid" "$(python3 dev/ci/check-mermaid.py --cli-spec)"
# puppeteer's postinstall fetches the browser mmdc drives, but it catches its own
# download failures and exits 0, so a flaky fetch leaves a green install step and an
# mmdc that cannot launch. Fetch it again here, where a failure fails the job.
npm --prefix "$RUNNER_TEMP/mermaid" exec -- puppeteer browsers install chrome-headless-shell
echo "$RUNNER_TEMP/mermaid/node_modules/.bin" >> "$GITHUB_PATH"
- name: Check mermaid diagrams render
run: python3 dev/ci/check-mermaid.py
- name: Check missing suites
run: python3 dev/ci/check-suites.py
- name: Check micro benchmark runner
run: python3 dev/ci/check-benchmark-runner.py
- name: Check pull request type labeling
run: node --test dev/ci/pr-type-label.test.mjs
- name: Check Iceberg shard inventory validation
run: python3 dev/ci/test-iceberg-shards.py
- name: Check Iceberg write report summary
run: python3 dev/ci/test-summarize-iceberg-writes.py
- name: Check CI config invariants
run: python3 dev/ci/check-ci-config.py
- name: Install actionlint
# Pure network, and preflight gates every other job, so a single reset
# connection here would fail the whole run. Download to a file rather
# than pipe into bash so a failed download cannot run a partial script.
run: |
for attempt in 1 2 3; do
if curl -sSfL --retry 3 --retry-all-errors -o download-actionlint.bash \
https://raw.githubusercontent.com/rhysd/actionlint/main/scripts/download-actionlint.bash \
&& bash download-actionlint.bash; then
break
fi
if [ "$attempt" = 3 ]; then
echo "::error::actionlint download failed after 3 attempts."
exit 1
fi
echo "::warning::actionlint download failed (attempt $attempt of 3); retrying in $((attempt * 10))s."
sleep $((attempt * 10))
done
echo "$PWD" >> "$GITHUB_PATH"
- name: Lint GitHub Actions workflows
run: actionlint -color --shellcheck=off
# ---------------------------------------------------------------------------
# changes: compute which long jobs need to run for this event. Replaces the
# per-workflow `on: paths:` filters that used to gate triggering. Filter
# rules live in dev/ci/compute-changes.py, which is invoked here in lieu of
# dorny/paths-filter (not on the apache org actions allow list). On
# workflow_dispatch every output is forced true so a manual run can
# exercise any gated job.
# ---------------------------------------------------------------------------
changes:
name: Detect changes
needs: preflight
runs-on: ubuntu-slim
# `actions: read` is for dev/ci/nightly-base.py, which lists previous
# scheduled runs to pick the nightly's diff base.
permissions:
actions: read
contents: read
outputs:
build_linux: ${{ steps.compute.outputs.build_linux }}
build_linux_full: ${{ steps.compute.outputs.build_linux_full }}
build_linux_all_profiles: ${{ steps.compute.outputs.build_linux_all_profiles }}
build_macos: ${{ steps.compute.outputs.build_macos }}
benchmark: ${{ steps.compute.outputs.benchmark }}
delta_gate: ${{ steps.compute.outputs.delta_gate }}
pyarrow_udf: ${{ steps.compute.outputs.pyarrow_udf }}
docs: ${{ steps.compute.outputs.docs }}
spark_3_4: ${{ steps.compute.outputs.spark_3_4 }}
spark_3_5: ${{ steps.compute.outputs.spark_3_5 }}
spark_4_0: ${{ steps.compute.outputs.spark_4_0 }}
spark_4_1: ${{ steps.compute.outputs.spark_4_1 }}
spark_4_1_hive: ${{ steps.compute.outputs.spark_4_1_hive }}
iceberg_1_8: ${{ steps.compute.outputs.iceberg_1_8 }}
iceberg_1_9: ${{ steps.compute.outputs.iceberg_1_9 }}
iceberg_1_10: ${{ steps.compute.outputs.iceberg_1_10 }}
iceberg_1_11: ${{ steps.compute.outputs.iceberg_1_11 }}
steps:
- uses: actions/checkout@v7
with:
# Need both branches' history so we can diff base..head for PRs and
# before..after for pushes.
fetch-depth: 0
- name: Compute outputs
id: compute
shell: bash
env:
EVENT_NAME: ${{ github.event_name }}
EVENT_ACTION: ${{ github.event.action }}
LABEL_NAME: ${{ github.event.label.name }}
PR_LABELS: ${{ toJSON(github.event.pull_request.labels.*.name) }}
PR_BASE_SHA: ${{ github.event.pull_request.base.sha }}
PR_HEAD_SHA: ${{ github.event.pull_request.head.sha }}
MQ_BASE_SHA: ${{ github.event.merge_group.base_sha }}
MQ_HEAD_SHA: ${{ github.event.merge_group.head_sha }}
PUSH_BEFORE: ${{ github.event.before }}
PUSH_AFTER: ${{ github.sha }}
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
set -euo pipefail
: > changed_files.txt
if [[ "$EVENT_NAME" == "workflow_dispatch" ]]; then
# No meaningful base to diff against; compute-changes.py forces
# every output true for this event so a manual run can exercise
# any gated job.
:
elif [[ "$EVENT_NAME" == "schedule" ]]; then
# The nightly's base is the commit the last successful scheduled
# run tested, so everything that landed since is covered exactly
# once and a red nightly keeps its commits in scope until a green
# one supersedes it. A quiet day diffs to nothing and runs
# nothing; so does a docs-only day.
#
# Without a base -- the first nightly, an API error, or a base no
# longer on main -- run the whole tier by treating every tracked
# file as changed, the same way the first push to a branch does
# below. Guessing a narrower base would be worse than running
# nothing: a window that starts after a commit no nightly has
# covered yet skips the suites that commit needs, goes green
# anyway, and then becomes the base for tomorrow, dropping that
# coverage for good.
prev=$(python3 dev/ci/nightly-base.py)
if [[ -n "$prev" ]] && git merge-base --is-ancestor "$prev" HEAD 2>/dev/null; then
echo "Nightly base: $prev"
git diff --name-only "$prev"..HEAD > changed_files.txt
else
echo "Nightly base unknown; running the whole nightly tier"
git ls-tree -r --name-only HEAD > changed_files.txt
fi
elif [[ "$EVENT_NAME" == "pull_request" ]]; then
git diff --name-only "$PR_BASE_SHA"..."$PR_HEAD_SHA" > changed_files.txt
elif [[ "$EVENT_NAME" == "merge_group" ]]; then
# The queue branch is built on top of base_sha, so this covers every
# entry batched into this merge group, not just the newest one.
git diff --name-only "$MQ_BASE_SHA"..."$MQ_HEAD_SHA" > changed_files.txt
else
# push to main; first push to a branch has all-zero before sha
if [[ "$PUSH_BEFORE" =~ ^0+$ ]]; then
git ls-tree -r --name-only "$PUSH_AFTER" > changed_files.txt
else
git diff --name-only "$PUSH_BEFORE".."$PUSH_AFTER" > changed_files.txt
fi
fi
echo "Changed files:"
cat changed_files.txt
python3 dev/ci/compute-changes.py changed_files.txt >> "$GITHUB_OUTPUT"
# ---------------------------------------------------------------------------
# Heavy jobs: each is a thin caller of an existing reusable workflow, gated
# on the one `changes` output that covers it.
#
# That single output already folds in everything these gates used to spell
# out inline: which files the job cares about, which events may run it, and
# which opt-in label it needs on a pull request. All of it lives in FILTERS
# and POLICY in dev/ci/compute-changes.py, where it is one table instead of
# ten near-identical `${{ }}` expressions, and where dev/ci/check-ci-config.py
# can actually test it.
# ---------------------------------------------------------------------------
pr_build_linux:
name: PR Build (Linux)
needs: changes
# Three POLICY outputs feed one call, the same shape as spark_4_1 below.
# `build_linux` decides whether the workflow runs at all; `build_linux_full`
# whether it runs the lints and the test matrix as well as the jobs that
# populate main's actions/cache entries; `build_linux_all_profiles` whether
# the test matrix covers every Spark profile or only the default one. The
# combinations that occur:
#
# pull request linux, full -> profiles: pr
# ... with the label linux, full, all -> profiles: all
# `labeled` run all -> profiles: nightly
# merge queue linux, full -> profiles: pr
# nightly all -> profiles: nightly
# push to main linux -> cache-refresh-only
#
# Only push to main sets `build_linux` without `build_linux_full`, which is
# the whole point: the queue has already tested that tree, so the push run
# is there for the caches alone. A `labeled` run and the nightly set only
# the third, and then run just the profiles the PR and queue tiers skip.
if: needs.changes.outputs.build_linux == 'true' || needs.changes.outputs.build_linux_all_profiles == 'true'
uses: ./.github/workflows/pr_build_linux.yml
with:
cache-refresh-only: ${{ needs.changes.outputs.build_linux_full != 'true' && needs.changes.outputs.build_linux_all_profiles != 'true' }}
profiles: >-
${{ needs.changes.outputs.build_linux_all_profiles != 'true' && 'pr'
|| needs.changes.outputs.build_linux_full != 'true' && 'nightly'
|| 'all' }}
pr_build_macos:
name: PR Build (macOS)
needs: changes
# Queue-only by default; PRs need the `run-macos-tests` label.
if: needs.changes.outputs.build_macos == 'true'
uses: ./.github/workflows/pr_build_macos.yml
pr_benchmark_check:
name: PR Benchmark Check
needs: changes
# Queue-only by default; PRs need the `run-benchmark-check` label.
if: needs.changes.outputs.benchmark == 'true'
uses: ./.github/workflows/pr_benchmark_check.yml
delta_build_gate:
name: Delta Contrib Build Gate
needs: changes
# Queue-only by default; PRs need the `run-delta-build-gate` label.
if: needs.changes.outputs.delta_gate == 'true'
# The called workflow needs only a checkout.
permissions:
contents: read
uses: ./.github/workflows/delta_build_gate.yml
pyarrow_udf_test:
name: PyArrow UDF Tests
needs: changes
# Queue-only by default; PRs need the `run-pyarrow-udf-tests` label.
if: needs.changes.outputs.pyarrow_udf == 'true'
# The called workflow needs only a checkout.
permissions:
contents: read
uses: ./.github/workflows/pyarrow_udf_test.yml
docs:
name: Deploy Comet site
needs: changes
# docs deploys to asf-site, so only run on push-to-main (or a manual dispatch).
if: needs.changes.outputs.docs == 'true'
uses: ./.github/workflows/docs.yaml
spark_3_4:
name: Spark SQL Tests (Spark 3.4)
needs: changes
# Spark 3.4 is deprecated, so this is the one test job the merge queue does
# not run: it needs the `run-spark-3.4-tests` label on a pull request, or a
# manual workflow_dispatch.
#
# It is still in `required_checks.needs`, but be precise about what that
# buys. Adding the label fires a `labeled` event, and on that event the
# aggregator deliberately publishes the advisory `Required Checks (label
# run)` name instead of the required one (see its comment below), so a red
# 3.4 there leaves the commit's existing verdict alone. 3.4 only counts
# toward the required `Required Checks` on a pull_request event that is not
# `labeled` -- a push with the label already applied. With no queue run left
# to catch it, that push is the only thing that turns a 3.4 failure into a
# merge blocker.
if: needs.changes.outputs.spark_3_4 == 'true'
uses: ./.github/workflows/spark_sql_test_reusable.yml
with:
spark-short: '3.4'
spark-full: '3.4.3'
java: 17
spark_3_5:
name: Spark SQL Tests (Spark 3.5)
needs: changes
# Nightly by default; PRs need the `run-spark-3.5-tests` label.
if: needs.changes.outputs.spark_3_5 == 'true'
uses: ./.github/workflows/spark_sql_test_reusable.yml
with:
spark-short: '3.5'
spark-full: '3.5.9'
java: 17
spark_4_0:
name: Spark SQL Tests (Spark 4.0)
needs: changes
# Nightly by default; PRs need the `run-spark-4.0-tests` label. Spark 4.1
# is the one Spark SQL suite the queue runs, being the default profile.
if: needs.changes.outputs.spark_4_0 == 'true'
uses: ./.github/workflows/spark_sql_test_reusable.yml
with:
spark-short: '4.0'
spark-full: '4.0.4'
java: 17
spark_4_1:
name: Spark SQL Tests (Spark 4.1)
needs: changes
# Queue-only by default, like every other Spark SQL suite. Two POLICY
# outputs feed one call, so the queue gets every module from a single
# 40-minute build instead of two. `spark_4_1` covers catalyst and the
# sql_core shards; `spark_4_1_hive` adds the sql_hive shards. On a pull
# request `run-spark-4.1-tests` sets both, and `run-spark-4.1-hive-tests`
# sets only the second, which then runs only the hive rows.
if: needs.changes.outputs.spark_4_1 == 'true' || needs.changes.outputs.spark_4_1_hive == 'true'
uses: ./.github/workflows/spark_sql_test_reusable.yml
with:
spark-short: '4.1'
spark-full: '4.1.3'
java: 17
modules: >-
${{ needs.changes.outputs.spark_4_1 != 'true' && 'hive'
|| needs.changes.outputs.spark_4_1_hive != 'true' && 'core'
|| 'all' }}
iceberg_1_8:
name: Iceberg Spark SQL Tests (Iceberg 1.8)
needs: changes
# Nightly by default; PRs need the `run-iceberg-tests` label.
if: needs.changes.outputs.iceberg_1_8 == 'true'
uses: ./.github/workflows/iceberg_spark_test_reusable.yml
with:
iceberg-short: '1.8'
iceberg-full: '1.8.1'
spark-short: '3.4'
spark-full: '3.4.3'
java: 17
iceberg_1_9:
name: Iceberg Spark SQL Tests (Iceberg 1.9)
needs: changes
# Nightly by default; PRs need the `run-iceberg-tests` label.
if: needs.changes.outputs.iceberg_1_9 == 'true'
uses: ./.github/workflows/iceberg_spark_test_reusable.yml
with:
iceberg-short: '1.9'
iceberg-full: '1.9.1'
spark-short: '3.5'
spark-full: '3.5.9'
java: 17
iceberg_1_10:
name: Iceberg Spark SQL Tests (Iceberg 1.10)
needs: changes
# Nightly by default; PRs need the `run-iceberg-tests` label.
if: needs.changes.outputs.iceberg_1_10 == 'true'
uses: ./.github/workflows/iceberg_spark_test_reusable.yml
with:
iceberg-short: '1.10'
iceberg-full: '1.10.0'
spark-short: '3.5'
spark-full: '3.5.9'
java: 17
iceberg_1_11:
name: Iceberg Spark SQL Tests (Iceberg 1.11)
needs: changes
# Queue-only by default; PRs need the `run-iceberg-tests` label. Iceberg
# 1.11 is our only Spark 4.1 Iceberg coverage, which is why it is the one
# Iceberg version the queue runs while 1.8/1.9/1.10 run nightly.
if: needs.changes.outputs.iceberg_1_11 == 'true'
uses: ./.github/workflows/iceberg_spark_test_reusable.yml
with:
iceberg-short: '1.11'
iceberg-full: '1.11.0'
spark-short: '4.1'
spark-full: '4.1.3'
java: 17
# ---------------------------------------------------------------------------
# required_checks: the single context listed in `required_status_checks` for
# `main` in `.asf.yaml`, and therefore the only thing the merge queue waits
# on.
#
# None of the jobs above can be required directly, because the name a caller
# of a reusable workflow publishes depends on whether it ran:
#
# skipped by `if:` one check run named exactly `PR Build (Linux)`
# actually ran only `PR Build (Linux) / Spark 4.1, JDK 17 [exec]`,
# ... and no bare `PR Build (Linux)` at all
#
# So requiring the bare name would block every code change, and requiring a
# nested name would block every docs-only change. Both hang rather than fail,
# and a required context that never reports also locks `.asf.yaml` itself,
# which then needs an INFRA ticket to unwedge. Aggregating into one flat job,
# whose name is published on every event, avoids the whole class of problem.
#
# `if: always()` is what makes this work: without it the job inherits the
# default `success()` and is itself skipped the moment any dependency fails.
#
# The name is an expression because ci_label.yml runs this file on `labeled`
# too, and on that event POLICY deliberately skips the PR tier (it already
# ran at this commit). A label run's verdict therefore says nothing about the
# commit's applicable suites, so publishing it as `Required Checks` could let
# a `dependencies` label turn a still-running or red commit run green (issue
# #5007). Label runs publish under a name nothing requires instead. Skipping
# the job here would not help: a skipped check run still carries the name and
# still counts as passing. Running label runs as a separate workflow (issue
# #6159) is the other half: it stops a label run from hiding this job's
# verdict from the merge box. dev/ci/check-ci-config.py enforces all of it.
# ---------------------------------------------------------------------------
required_checks:
name: ${{ github.event.action == 'labeled' && 'Required Checks (label run)' || 'Required Checks' }}
if: always()
# Reads only the `needs` context, so it needs no token access at all.
permissions: {}
needs:
- preflight
- changes
- pr_build_linux
- pr_build_macos
- pr_benchmark_check
- delta_build_gate
- pyarrow_udf_test
- spark_3_4
- spark_3_5
- spark_4_0
- spark_4_1
- iceberg_1_8
- iceberg_1_9
- iceberg_1_10
- iceberg_1_11
runs-on: ubuntu-slim
steps:
- name: Summarize upstream results
env:
NEEDS: ${{ toJSON(needs) }}
run: echo "$NEEDS"
# `skipped` is a pass: it means the change did not touch anything that
# job covers. Only `failure` and `cancelled` block the merge.
- name: Fail if any upstream job failed or was cancelled
if: contains(needs.*.result, 'failure') || contains(needs.*.result, 'cancelled')
run: |
echo "::error::One or more upstream jobs did not succeed. See the results above."
exit 1
# ---------------------------------------------------------------------------
# nightly_report: a red nightly has no pull request to show up on, and a
# scheduled run only emails whoever last touched the workflow file, so
# without this a regression in the nightly tier sits unnoticed on the
# Actions tab. On the scheduled event, and only when `required_checks` is
# not green, it opens an issue labelled `ci-nightly-failure` listing the
# failed jobs, or comments on the one that is already open so consecutive
# red nights accumulate in one place instead of one issue a day. Closing
# the issue is how the failure is acknowledged; the next red night opens a
# new one.
#
# It sits downstream of the aggregator rather than of the nightly jobs
# themselves so the job list here does not have to be kept in step with
# POLICY.
# ---------------------------------------------------------------------------
nightly_report:
name: Nightly failure report
needs: required_checks
if: always() && github.event_name == 'schedule' && needs.required_checks.result != 'success'
permissions:
actions: read
issues: write
runs-on: ubuntu-slim
steps:
- uses: actions/github-script@v9
with:
script: |
const label = 'ci-nightly-failure';
const { owner, repo } = context.repo;
const runUrl = `${context.serverUrl}/${owner}/${repo}/actions/runs/${context.runId}`;
const shortSha = context.sha.slice(0, 10);
const day = new Date().toISOString().slice(0, 10);
const jobs = await github.paginate(github.rest.actions.listJobsForWorkflowRun, {
owner, repo, run_id: context.runId, filter: 'latest', per_page: 100,
});
const failed = jobs
.filter((job) => job.conclusion === 'failure' || job.conclusion === 'cancelled')
.map((job) => `- [${job.name}](${job.html_url}) (${job.conclusion})`);
const body = [
`The nightly CI run against \`main\` at ${shortSha} did not pass: ${runUrl}`,
'',
failed.length ? 'Jobs that did not succeed:' : 'No individual job reported a failure; see the run for details.',
...failed,
'',
'The nightly tier runs the Spark and Iceberg versions the merge queue does not',
'(see `POLICY` in `dev/ci/compute-changes.py`). Close this issue once the failure',
'is understood; the next red nightly opens a new one.',
].join('\n');
const { data: open } = await github.rest.issues.listForRepo({
owner, repo, labels: label, state: 'open', per_page: 1,
});
if (open.length > 0) {
await github.rest.issues.createComment({ owner, repo, issue_number: open[0].number, body });
core.info(`Commented on #${open[0].number}`);
} else {
const { data: issue } = await github.rest.issues.create({
owner, repo, title: `Nightly CI failed on ${day}`, body, labels: [label],
});
core.info(`Opened #${issue.number}`);
}