Repository navigation
Conversation
- Model registry: training scripts now only register new versions; champion promotion is a separate gated decision (promote_champion_if_better) exposed via scripts/promote_model.py and make promote, wired into run_training.sh - Retrain DAG: now actually runnable in Airflow. ./training volume-mounted read-only, training deps installed in the image (apache/airflow:2.7.3-python3.11 base to match the training venv), gate task delegates to the shared promotion function, subprocess output captured into task logs, manual trigger (schedule=None) with a production note. Verified end-to-end via live DAG runs - MLflow: proxied artifact storage (--serve-artifacts, --artifacts-destination, --default-artifact-root mlflow-artifacts:/) so clients upload/download over HTTP instead of writing artifacts to their own filesystem - Monitoring: HighFraudRate alert wrapped both ratio sides in sum() — PromQL label matching made the old expression divide each series by itself; validated with promtool - Cleanup: removed phantom-phase leftovers (Kafka/Zookeeper compose stubs, test_kafka_flow.py, root tests/conftest.py stub, evidently drift_report.py stub); README quickstart fixed (admin/admin creds, restart serving after training) - Docs: wiki, service READMEs, and code references updated to match; AUDIT.md added with all findings and resolution notes
- Features: drop leaky/dead amount_zscore everywhere (32-column set); drift generator is_night now matches the canonical 22:00-06:00 definition (L6) - Training: 60/20/20 split, threshold tuned on val, metrics reported on held-out test (M3); SMOTE as the single imbalance strategy, scale_pos_weight and use_label_encoder removed (M2, L4); both models retrained on the new schema - Serving: Time required in the API schema (M4); dead prepare_features_batch deleted (M5); autoencoder load_context normalises Windows artifact-path separators so Windows-logged models load on Linux - Registry: promote_model.py --force for migrations where gate metrics are not comparable; cp1252-safe ASCII arrows in prints - Docker: serving/.dockerignore keeps .venv, tests, and caches out of the build context; image 12.1 GB -> 9.9 GB (M6) - CI: typecheck job is now blocking and training tests run in CI (M7) - Docs: wiki, README, plan.md, EDA notebook, and AUDIT.md resolution notes synced with all of the above
- Airflow: validate_features checks the parquet schema via pyarrow before reading data, making the column check reachable (L2) - Serving: shap pinned to the deployed 0.49.1 (L3); error counter uses the same full model_name labels as the other metrics (L5); scalers receive DataFrames instead of bare arrays, silencing the sklearn feature-names warning (L8); autoencoder retrained as v5 to bake the pyfunc-side fix into the registry - Training: find_optimal_threshold rewritten as a vectorised cumulative-sum sweep, equivalence-tested against the loop on ~900 randomised cases including score ties (L10) - Compose: healthchecks for mlflow and serving; serving depends on mlflow service_healthy (removes the model-load startup race), and its own healthcheck reports healthy only when both models are loaded (L11) - Notebook: eda.ipynb committed executed so charts render on GitHub (L7) - Docs: plan.md describes the implemented degraded-mode behaviour (L9); wiki 01-06 synced; AUDIT.md closes with all findings resolved
…matting - Training: torch.manual_seed(RANDOM_STATE) in the autoencoder so weight init and shuffling are reproducible; verified with two consecutive runs producing identical metrics (AUC-ROC 0.9601, PR-AUC 0.2998, F1 0.4967) - Serving: protected_namespaces cleared on Settings, PredictionResponse, and BatchPredictionItem, silencing six pydantic warnings on import; the one remaining deprecation warning is inside mlflow 2.9 itself - Tooling: black[jupyter] in requirements-dev.txt and the CI lint job so the notebook is format-checked consistently; eda.ipynb reformatted - Docs: wiki 03 venv-split rationale reworded (isolation and pinning, not image size — serving ships torch for the in-process pyfunc); wiki 02/05/06 line references synced
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.