Auditable human-in-the-loop denial review for document-heavy healthcare ops
π Live demo: https://workbench.yuanliu.dev
Try case-002 for the main demo arc (missing documents path and approval flow).
Denial review pipelines process a full case β extraction, gap analysis, draft β typically in under 15 seconds. Replay any approved run: status is preserved, no model call is made, and the JSONL log is unchanged. Verify it yourself: ./bench.sh (pipeline), ./scripts/replay_integrity.sh (replay invariant), and ./scripts/export_to_vifei.sh (share-safe Vifei export)
Denial Review Workbench is a thin, operator-ready workflow tool for reviewing insurance authorization denials. It ingests a case packet β denial letter, auth request, clinical notes β runs a structured AI pipeline to extract facts, identify missing documentation, and draft a next action, then routes the result to a human reviewer for approval or escalation. Every decision is logged to an append-only JSONL event store and can be replayed deterministically without touching the model.
A caseworker opens case-002 β a skilled nursing authorization denied for insufficient documentation. The left panel loads the denial letter, auth request, clinical notes, and the matching policy excerpt. They click Run Review.
The pipeline extracts facts from the denial packet, retrieves the relevant policy sections by keyword match, identifies two missing items β signed physician order and progress notes within 30 days β and drafts a next action. Confidence renders as a color-banded percentage. Each evidence reference is clickable and highlights the source quote in the document panel.
The reviewer edits one line of the draft, clicks Approve, and the run transitions to approved. The approved event carries final_recommendation β exactly what the human signed off on β as a first-class audit artifact.
Opening Run History shows the full event timeline. Clicking Replay reconstructs the complete run state from the JSONL log without calling the model. An amber REPLAY badge confirms the reconstruction is read-only.
case-003 demonstrates the escalation branch: conflicting denial reasons trigger should_escalate: true, the draft and approve UI are suppressed, and the center panel shows an escalation notice. The server returns 409 on any approve attempt β the guard is enforced at the API layer, not just the UI.
The triage queue lists all cases with facility, scenario, expected path, and summary columns. Reviewers pick a case and click Run Review to start the pipeline.
A persistent summary header shows status, facility, confidence, missing-item count, and start time. Left panel: denial letter, auth request, clinical notes, and retrieved policy excerpts. Center panel: extracted facts with color-banded confidence, gap analysis table, and editable recommendation. Right panel: event timeline.
Clicking an evidence reference scrolls to the cited document and highlights the source quote with a yellow marker. An "Evidence focus" badge on the document confirms which source is active.
The reviewer edits the draft text and clicks Approve. The status transitions to approved, the textarea locks into a read-only "Final Text" display, and the event timeline gains an "Approved" entry.
Clicking Replay reconstructs the full run state from the JSONL event log β no model call is made. An amber REPLAY badge confirms the view is a read-only reconstruction.
Conflicting denial reasons trigger should_escalate: true. A red banner blocks the workflow β no draft is generated, no Approve button is shown. The server enforces this with a 409 guard on the approve endpoint.
flowchart TD
UI["Three-Panel UI<br/>Documents + Policy Β· Facts + Gap + Draft Β· History"]
POST["POST /runs (sync)"]
subgraph BACKEND["FastAPI Backend<br/>At most 3 model calls per run"]
EF["extract_facts<br/>model call 1"]
PS["policy_search"]
AG["analyze_gap<br/>model call 2"]
DN["draft_next_action<br/>model call 3<br/>(if not escalated)"]
AC["anthropic_client.py<br/>single entry for model calls"]
EF --> PS --> AG --> DN
AC -.-> EF
AC -.-> AG
AC -.-> DN
end
subgraph STORE["JSONL Event Store"]
JSONL["data/runs/{run_id}.jsonl<br/>Append-only. Never rewritten.<br/>Full payloads for replay."]
end
subgraph REPLAY["replay.py"]
R["Read-only JSONL reconstruction<br/>Never calls model providers<br/>Never writes new events"]
end
UI --> POST --> EF
EF --> JSONL
AG --> JSONL
DN --> JSONL
JSONL --> R
Three model calls, strict order:
extract_factsβ payer, service, denial reason, confidence, evidence refsanalyze_gapβ missing items, conflicts, escalation decisiondraft_next_actionβ only whenshould_escalateis false
V1 run statuses: approval_requested β approved or escalated
Replay invariant: The approved event carries final_recommendation β exactly what the human signed off on. Replay reconstructs this from the JSONL log without any model call.
Run logs are append-only JSONL. Each event carries a type, ISO 8601 timestamp, and full payload sufficient for deterministic replay.
Replay integrity proof: scripts/replay_integrity.sh replays an approved run 10 times and verifies that every response is identical, is_replay_response is true, and the JSONL log gains zero lines. Run it against a live backend:
./scripts/replay_integrity.shReplay 1: completed in 48ms, JSONL lines: 3 (unchanged) β
Replay 2: completed in 50ms, JSONL lines: 3 (unchanged) β
Replay 3: completed in 46ms, JSONL lines: 3 (unchanged) β
Replay 4: completed in 42ms, JSONL lines: 3 (unchanged) β
Replay 5: completed in 46ms, JSONL lines: 3 (unchanged) β
Replay 6: completed in 51ms, JSONL lines: 3 (unchanged) β
Replay 7: completed in 48ms, JSONL lines: 3 (unchanged) β
Replay 8: completed in 43ms, JSONL lines: 3 (unchanged) β
Replay 9: completed in 58ms, JSONL lines: 3 (unchanged) β
Replay 10: completed in 46ms, JSONL lines: 3 (unchanged) β
10/10 replays verified. Zero events written during replay.
Vifei export proof: scripts/export_to_vifei.sh normalizes a workbench run into Vifei's CommittedEvent format and produces a share-safe bundle via vifei export. The workbench JSONL is not Vifei EventLog β the script bridges the two schemas without modifying the source log. The bridge contract is documented in docs/vifei-bridge-contract.md.
./scripts/export_to_vifei.sh run-case-002-20260416075553Normalized 8 events β out/run-case-002-20260416075553.vifei.jsonl
Running: vifei export out/run-case-002-20260416075553.vifei.jsonl --share-safe --output out/run-case-002-20260416075553.tar.zst
{"code":"OK","data":{"blob_count":0,"bundle_hash":"20f96940ae97c23ad7b74f2f3793c0c775bb58ca29af7c418399d49ebfa3d569","event_count":8},"exit_code":0,"message":"Export completed successfully.","ok":true,"schema_version":"vifei-cli-robot-v1.1"}
Export succeeded.
Bundle: out/run-case-002-20260416075553.tar.zst
Bundle size: 2131 bytes
The bundle is a deterministic tar.zst archive with a BLAKE3 hash. Escalated runs also export (5 source events normalize to 6 with a synthetic RunEnd). Mock and partial runs are rejected with explicit messages. If the vifei binary is not installed, the script still produces the normalized JSONL and prints build instructions.
| Layer | Technology |
|---|---|
| Backend | FastAPI, Python 3.11+, Pydantic v2, uvicorn |
| Frontend | React 18, TypeScript, Vite, Bun |
| Model | Anthropic claude-sonnet-4-6 |
| Storage | Append-only JSONL event log |
| Schemas | Pydantic (backend) mirrored to TypeScript interfaces (frontend) |
Prerequisites: Python 3.11+, Bun, Anthropic API key
git clone https://github.com/yuan-cloud/denial-review-workbench
cd denial-review-workbenchBackend:
cp backend/.env.example backend/.env
# add your Anthropic API key to backend/.env
cd backend
# Load .env: bash/zsh: source .env
# fish: export (cat .env | psub)
pip install fastapi uvicorn anthropic pydantic
uvicorn app.main:app --reload --port 8000Verify: curl http://localhost:8000/health
Expected: {"status":"ok","phase":"3"}
Frontend:
cd frontend
bun install
bun run devOpens at http://localhost:5274
Environment variables:
| Variable | Default | Description |
|---|---|---|
ANTHROPIC_API_KEY |
β | Required for live pipeline |
VITE_PORT |
5274 |
Frontend port |
FRONTEND_URL |
http://localhost:5274 |
CORS allowed origin |
VITE_API_BASE |
http://localhost:8000 |
Backend URL for frontend |
Three synthetic cases cover the full decision space:
| Case | Path | Expected outcome |
|---|---|---|
case-001 |
Happy path β all docs present | approve_or_proceed, confidence β₯ 0.85 |
case-002 |
Missing docs β main demo case | request_missing_documents, 2 missing items |
case-003 |
Conflicting denial reasons | escalated, recommendation null |
Policy files live in data/policies/. Swapping policy packs and approval rules is how this scales across facilities β the orchestration spine stays the same.
Seeded mock runs vs. live runs: On startup, the backend loads mock_run.json from each case directory into in-memory state (IDs like run-case-002-mock). These are fixed snapshots for offline exploration. A live demo run (./bench.sh case-002) produces a fresh timestamped run ID with a complete JSONL audit trail.
If a run fails, use the case-002 fallback endpoint:
GET /demo-fallback/case-002
Returns the saved mock run for case-002 so you can inspect the Recommendation or Run History panel state.
Built with multiple AI coding agents (Claude Code, Codex) working in parallel from a shared operating contract (AGENTS.md). Agents coordinate through message passing and file reservations, commit independently to main, and run automated bug scanning before every commit. Task graph, session search, and orchestration are scripted β no manual dispatch.
# Use absolute path β some environments alias grep to
# ripgrep (rg), which uses different flag syntax and
# silently fails with --include flags
/usr/bin/grep -rE "PHI|patient_name|ssn|date_of_birth" \
data/ backend/ frontend/src/ \
--include="*.py" --include="*.tsx" \
--include="*.json" --include="*.md"Returns empty. No patient identifiers in any synthetic case data.





