Skip to content

[pull] main from hookdeck:main#160

Merged
pull[bot] merged 2 commits into
erickirt:mainfrom
hookdeck:main
Jun 25, 2026
Merged

[pull] main from hookdeck:main#160
pull[bot] merged 2 commits into
erickirt:mainfrom
hookdeck:main

Conversation

@pull

@pull pull Bot commented Jun 25, 2026

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

leggetter and others added 2 commits June 25, 2026 09:59
…977)

* Redact eval artifact secrets and clarify quickstart polling snippets.

Address Copilot review on #976: redact known secrets when writing
transcripts and judge failures, and scan results/runs before CI upload.

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* Add tests for eval artifact secret redaction.

Cover pattern and literal env redaction, JSON artifact output, and wire
npm run test:redact-secrets into npm run test.

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* Extend eval artifact redaction to webhook URLs and CI env.

Add EVAL_TEST_DESTINATION_URL and OUTPOST_TEST_WEBHOOK_URL to literal
redaction; pass secrets into the CI redact step so re-scan can replace
plain JSON echoes (Copilot review on #977).

Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix eval artifact redaction corrupting JSON structure.

Deep-walk artifact objects and redact string leaves before serialization so
transcript.json stays parseable for heuristic scoring.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* Reconcile LLM judge overall pass with criteria array.

Derive overall_transcript_pass from criteria when present so eval-ci does
not fail when the model marks every criterion pass but overall false.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Harden LLM judge boolean parsing per review feedback.

Parse string "true"/"false" explicitly, apply the same to criteria pass
fields, only log reconciliation when overall was explicitly set, and align
test assertion messages with other eval unit tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
@pull pull Bot locked and limited conversation to collaborators Jun 25, 2026
@pull pull Bot added the ⤵️ pull label Jun 25, 2026
@pull
pull Bot merged commit 7562289 into erickirt:main Jun 25, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant