Automated integration-failure triage for the daily CI window 2026-09-10 → 2026-09-16.
This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.
Serge dispatched one task per failure group below — each opens or updates its own PR on a serge/fix/itf-<fingerprint> branch. This table is refreshed in place as Serge runs: a group links its #<pr> when opened, shows 🚫 no fix when Serge found no safe change, ⚠️ task failed on error, or (pending) while still running (a late PR links on the next nightly run).
Dispatched failure groups
| Model |
Error |
Occurrences |
PR |
qwen3_omni_moe (regressed by PR #48714) |
mixed — other (2) |
2 |
⚠️ task failed |
generation |
cuda_runtime — other (12) |
12 |
🚫 no fix |
deepseek_vl |
other — other (3) |
3 |
#48536 |
cwm |
OOM — other (2) |
2 |
🚫 no fix |
eomt_dinov3 |
output_mismatch — tensor values differ (2) |
2 |
#48882 |
Not dispatched — environment / dependency
These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).
2 models ran out of device memory (7 failures) — needs runner capacity, not a patch, so none of these were dispatched: bamba (4), llama4 (3).
5 group(s) skipped — stopped failing before this run (2 integration tests regressed by commit 774a675d827e (PR #47827), edgetam, depth_anything, glm4v_moe +1 more). Not dispatched and not awaiting a human.
Outcome recap
Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.
| Group |
Reason |
LLM |
Tokens (in / out) |
Failing tests |
qwen3_omni_moe (regressed by PR #48714) |
task runner exited without reporting (exit code 1) |
moonshotai/Kimi-K2.7-Code |
2,092,722 / 3,697 |
test_small_model_integration_test_batch_audio_matches_single · test_small_model_integration_test_w_audio |
generation |
[not_reproduced] GPU reproduce: the targeted tests did NOT fail at the base commit (not_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR. |
moonshotai/Kimi-K2.7-Code |
— / — |
test_validate_assistant · test_matches_immediate_check_at_max_length · test_matches_immediate_check_on_a_batch · +3 more |
cwm |
[not_reproduced] GPU reproduce: the targeted tests did NOT fail at the base commit (not_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR. |
moonshotai/Kimi-K2.7-Code |
— / — |
test_cwm_generation_20_tokens · test_cwm_sliding_window_long_sequence |
Generated 2026-09-17T00:12:11+00:00.
Automated integration-failure triage for the daily CI window
2026-09-10 → 2026-09-16.This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.
Serge dispatched one task per failure group below — each opens or updates its own PR on a
serge/fix/itf-<fingerprint>branch. This table is refreshed in place as Serge runs: a group links its#<pr>when opened, shows🚫 no fixwhen Serge found no safe change,⚠️ task failedon error, or(pending)while still running (a late PR links on the next nightly run).Dispatched failure groups
qwen3_omni_moe(regressed by PR #48714)generationdeepseek_vlcwmeomt_dinov3Not dispatched — environment / dependency
These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).
2 models ran out of device memory (7 failures) — needs runner capacity, not a patch, so none of these were dispatched:
bamba(4),llama4(3).5 group(s) skipped — stopped failing before this run (
2 integration tests regressed by commit 774a675d827e (PR #47827),edgetam,depth_anything,glm4v_moe+1 more). Not dispatched and not awaiting a human.Outcome recap
Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.
qwen3_omni_moe(regressed by PR #48714)moonshotai/Kimi-K2.7-Codegenerationnot_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR.moonshotai/Kimi-K2.7-Codecwmnot_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR.moonshotai/Kimi-K2.7-CodeGenerated 2026-09-17T00:12:11+00:00.