Skip to content

Initial Development phase completed - #1

Merged
Liuck27 merged 8 commits into
mainfrom
dev
Apr 22, 2026
Merged

Liuck27 merged 8 commits into
mainfrom
dev

Conversation

@Liuck27

@Liuck27 Liuck27 commented Apr 22, 2026

Copy link
Copy Markdown
Owner

No description provided.

Liuck27 added 8 commits April 12, 2026 19:46
Completes Phase 1 of the ML Fraud Detection Platform build-out:

Infrastructure
- docker-compose.yml: PostgreSQL 15 with healthcheck, named volume/network;
  all future services (MLflow, Airflow, Kafka, FastAPI, Prometheus, Grafana)
  present as commented stubs — each phase just uncomments
- scripts/init_db.sql: idempotently creates mlflow_db and airflow_db alongside
  fraud_db on first container start
- .env.example: all variables for all 11 phases documented with inline notes
- Makefile: full developer interface (up/down, psql, venv-*, lint, test, audit,
  go-build, check); Windows/Linux compatible via OS detection

Config & tooling
- .gitignore: covers Python artifacts, **/.venv/, data files, ML artifacts
  (*.pkl, *.pt, mlruns/), Go binaries, .env, .claude/
- .gitattributes: LF enforcement for Makefile, *.sh, *.sql, Dockerfile
- requirements-dev.txt: ruff, black, mypy, pytest, pip-audit

Per-service requirements stubs
- training, serving, airflow, streaming/producer, monitoring/evidently

Full directory skeleton with stub source files
- airflow/dags, airflow/plugins, feature_store, monitoring (Prometheus,
  Grafana provisioning, Evidently, alerting), serving/app (routes, models,
  schemas, tests), streaming (producer + Go consumer), training, tests/integration

plan.md updated
- Architecture diagram rewritten as Mermaid flowchart TD
- GitHub Actions CI added as Phase 10 (old Phase 10 renumbered to 11)
- Airflow: implement data_ingestion DAG (validate_csv → engineer_and_write),
  add Dockerfile using Airflow constraints, uncomment Airflow services in
  docker-compose.yml, update requirements (drop Feast, add pandas/pyarrow)
- Feature engineering: implement pure functions in airflow/plugins
  (log_transform_amount, extract_time_features, compute_interaction_features)
- Data download: implement scripts/download_data.py with Kaggle API integration
- EDA: add training/notebooks/eda.ipynb (class balance, amount/time
  distributions, V-feature discrimination, correlation, engineered features)
- Cleanup: remove out-of-scope streaming/, feature_store/, seed_data.py,
  train_isolation_forest.py; strip Kafka/Feast/Go from Makefile and .env.example
- Config: fix airflow-init restart policy and YAML command quoting in
  docker-compose.yml; generate valid Fernet key placeholder in .env.example
- Infrastructure: Dockerfile.mlflow (python:3.11-slim + mlflow + psycopg2); MLflow service uncommented in docker-compose.yml backed by PostgreSQL
- XGBoost: SMOTE oversampling + scale_pos_weight, StandardScaler, AUC-ROC 0.98, registered as fraud-xgboost@champion
- Autoencoder: PyTorch encoder-decoder trained on non-fraud only, TorchScript export, custom pyfunc wrapper, registered as fraud-autoencoder@challenger
- Shared utils: evaluate.py (compute_metrics, find_optimal_threshold, plot_roc/pr_curve); model_registry.py (promote_to_champion/challenger)
- Tooling: scripts/run_training.sh (cross-platform venv detection, set -euo pipefail); Makefile train/train-xgboost/train-autoencoder targets using PYTHON_TRAINING; pytest + mypy added to training/requirements.txt
…Prometheus

- Serving layer: config.py (pydantic-settings), schemas.py (full request/response models),
  loader.py (ModelRegistry: loads XGBoost + scaler + autoencoder pyfunc from MLflow),
  ab_testing.py (deterministic MD5-hash routing), explainer.py (SHAP TreeExplainer)
- Routes: GET /health, POST /predict (with SHAP), POST /predict/batch, GET /models,
  GET /metrics (Prometheus via prometheus-fastapi-instrumentator)
- Training fix: train_xgboost.py now logs StandardScaler as MLflow artifact so serving
  can apply identical scaling at inference time
- Infrastructure: serving/Dockerfile, package __init__.py files, pytest.ini warning filter;
  docker-compose.yml serving block uncommented
- Tests: 14 unit tests passing (5 A/B routing, 9 endpoint tests) with mocked MLflow
…ction

- Metrics: serving/app/metrics.py — 4 custom Prometheus metrics
  (inference_latency_seconds, inference_total, inference_errors_total,
  ab_test_assignments_total); predict.py instrumented for both /predict
  and /predict/batch
- Alerting: monitoring/alerting/rules.yml — HighFraudRate, HighInferenceLatency,
  InferenceErrorSpike rules
- Prometheus: monitoring/prometheus/prometheus.yml — fraud_serving scrape job
  targeting serving:8000/metrics
- Grafana: 4-panel dashboard (request rate, fraud %, p99 latency, A/B split)
  with auto-provisioning via dashboard.yml + datasource.yml (uid: prometheus-local)
- Drift: scripts/drift_report.py — Evidently DataDriftPreset, HTML → data/reports/;
  evidently requirements bumped to >=0.4.30,<0.5 for pydantic v2 compat
- Infrastructure: Prometheus + Grafana services uncommented in docker-compose.yml;
  Makefile gains up-monitoring and drift-report targets
- CI: GitHub Actions workflow with lint (ruff + black), typecheck (mypy), and unit test jobs; test job depends on lint for fast-fail
- Tests: 3 end-to-end integration tests (POST /predict schema, Prometheus counters, MLflow champion alias); conftest auto-skips when services are down; 7 new training unit tests for evaluate.py
- Airflow: retrain_dag.py implemented (validate features → train XGBoost → promote to champion if PR-AUC improved, @Weekly schedule)
- README: full rewrite with Mermaid architecture diagram, ≤10-command quickstart, key features, API reference, design decisions
- Fixes: black formatting across 11 files; ruff unused imports + E402 suppression for sys.path pattern; Makefile typecheck uses root venv mypy; serving/requirements.txt ML deps pinned to match training (xgboost==2.0.3, scikit-learn==1.3.2, etc.); docker-compose adds --serve-artifacts to MLflow and shares mlflow_artifacts volume with serving; train_autoencoder.py uses Path.as_posix() to avoid Windows backslash paths in MLmodel
- Docs: 10-part explanation wiki in docs/explanation/ covering big picture, ML concepts, infrastructure, data pipeline, training, serving API, monitoring, testing/CI, glossary, and extensions
- Docs: docs/explanation/README.md as wiki index with reading order and navigation
- Infra: LICENSE file added to repo root
- README: linked to wiki in introductory callout; minor copy updates
- plan.md: minor updates to match current project state
explainer.explain() was typed as list[dict[str, float | str]], causing
mypy to reject FeatureContribution(**c) in predict.py because the union
value type didn't match the specific str/float fields. ContributionDict
gives mypy exact per-key types so the **-unpacking is accepted.
@Liuck27
Liuck27 merged commit bf91a4a into main Apr 22, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant