Skip to content

Address audit findings and improve training, serving, and documentation - #3

Merged
Liuck27 merged 4 commits into
mainfrom
dev
Jun 12, 2026
Merged

Liuck27 merged 4 commits into
mainfrom
dev

Conversation

@Liuck27

@Liuck27 Liuck27 commented Jun 11, 2026

Copy link
Copy Markdown
Owner

No description provided.

Liuck27 added 4 commits June 10, 2026 23:39
- Model registry: training scripts now only register new versions; champion
  promotion is a separate gated decision (promote_champion_if_better) exposed
  via scripts/promote_model.py and make promote, wired into run_training.sh
- Retrain DAG: now actually runnable in Airflow. ./training volume-mounted
  read-only, training deps installed in the image (apache/airflow:2.7.3-python3.11
  base to match the training venv), gate task delegates to the shared promotion
  function, subprocess output captured into task logs, manual trigger
  (schedule=None) with a production note. Verified end-to-end via live DAG runs
- MLflow: proxied artifact storage (--serve-artifacts, --artifacts-destination,
  --default-artifact-root mlflow-artifacts:/) so clients upload/download over
  HTTP instead of writing artifacts to their own filesystem
- Monitoring: HighFraudRate alert wrapped both ratio sides in sum() — PromQL
  label matching made the old expression divide each series by itself;
  validated with promtool
- Cleanup: removed phantom-phase leftovers (Kafka/Zookeeper compose stubs,
  test_kafka_flow.py, root tests/conftest.py stub, evidently drift_report.py
  stub); README quickstart fixed (admin/admin creds, restart serving after
  training)
- Docs: wiki, service READMEs, and code references updated to match; AUDIT.md
  added with all findings and resolution notes
- Features: drop leaky/dead amount_zscore everywhere (32-column set);
  drift generator is_night now matches the canonical 22:00-06:00
  definition (L6)
- Training: 60/20/20 split, threshold tuned on val, metrics reported
  on held-out test (M3); SMOTE as the single imbalance strategy,
  scale_pos_weight and use_label_encoder removed (M2, L4); both
  models retrained on the new schema
- Serving: Time required in the API schema (M4); dead
  prepare_features_batch deleted (M5); autoencoder load_context
  normalises Windows artifact-path separators so Windows-logged
  models load on Linux
- Registry: promote_model.py --force for migrations where gate
  metrics are not comparable; cp1252-safe ASCII arrows in prints
- Docker: serving/.dockerignore keeps .venv, tests, and caches out
  of the build context; image 12.1 GB -> 9.9 GB (M6)
- CI: typecheck job is now blocking and training tests run in CI (M7)
- Docs: wiki, README, plan.md, EDA notebook, and AUDIT.md resolution
  notes synced with all of the above
- Airflow: validate_features checks the parquet schema via pyarrow
  before reading data, making the column check reachable (L2)
- Serving: shap pinned to the deployed 0.49.1 (L3); error counter
  uses the same full model_name labels as the other metrics (L5);
  scalers receive DataFrames instead of bare arrays, silencing the
  sklearn feature-names warning (L8); autoencoder retrained as v5 to
  bake the pyfunc-side fix into the registry
- Training: find_optimal_threshold rewritten as a vectorised
  cumulative-sum sweep, equivalence-tested against the loop on ~900
  randomised cases including score ties (L10)
- Compose: healthchecks for mlflow and serving; serving depends on
  mlflow service_healthy (removes the model-load startup race), and
  its own healthcheck reports healthy only when both models are
  loaded (L11)
- Notebook: eda.ipynb committed executed so charts render on GitHub (L7)
- Docs: plan.md describes the implemented degraded-mode behaviour
  (L9); wiki 01-06 synced; AUDIT.md closes with all findings resolved
…matting

- Training: torch.manual_seed(RANDOM_STATE) in the autoencoder so weight
  init and shuffling are reproducible; verified with two consecutive runs
  producing identical metrics (AUC-ROC 0.9601, PR-AUC 0.2998, F1 0.4967)
- Serving: protected_namespaces cleared on Settings, PredictionResponse,
  and BatchPredictionItem, silencing six pydantic warnings on import; the
  one remaining deprecation warning is inside mlflow 2.9 itself
- Tooling: black[jupyter] in requirements-dev.txt and the CI lint job so
  the notebook is format-checked consistently; eda.ipynb reformatted
- Docs: wiki 03 venv-split rationale reworded (isolation and pinning, not
  image size — serving ships torch for the in-process pyfunc); wiki 02/05/06
  line references synced
@Liuck27 Liuck27 self-assigned this Jun 11, 2026
@Liuck27
Liuck27 merged commit 86381f5 into main Jun 12, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant