Skip to content

Integrate managed benchmark service - #2

Merged
pvennacherla merged 7 commits into
mainfrom
integrate/managed-benchmark-service
Sep 30, 2026
Merged

pvennacherla merged 7 commits into
mainfrom
integrate/managed-benchmark-service

Conversation

@pvennacherla

Copy link
Copy Markdown
Collaborator

Summary

  • integrate the complete managed benchmark API, dashboard, artifact lifecycle, and DigitalOcean inference support
  • add Deep SWE, SWE-bench Verified, Terminal-Bench 2.1, and SWE Atlas QA/TW/RF support
  • preserve the DigitalOcean fork’s newer upstream changes and media tooling
  • include API documentation, retry workflows, request observability, and timeout handling

Validation

  • bun run format:check
  • bun run check
  • bun run typecheck
  • bun test (1,607 passed)
  • bun run build

jonathandieu and others added 7 commits August 10, 2026 17:19
…onventions

normalizeBaseUrl() unconditionally force-appended "/api/v1" to any custom
base URL, so pointing this harness at an OpenAI-compatible endpoint mounted
under a different path (e.g. DigitalOcean's inference-proxy at "/v1") always
404'd. It now only trims trailing slashes; callers pass a fully-qualified
base URL and it's used verbatim.

Also switch chat completions to streaming (stream: true, with
stream_options.include_usage). Long, high-reasoning-effort completions can
sit fully idle on a non-streaming request for many minutes and get killed by
network middleware before the origin finishes -- the same "long requests"
issue documented in Anthropic's own SDK
(https://github.com/anthropics/anthropic-sdk-python?tab=readme-ov-file#long-requests),
which is why inspect_ai auto-streams for Anthropic but not for this
OpenAI-compatible path. The response is reassembled into the same shape a
non-streaming response would have (falling back to the old plain-JSON path
when a response isn't SSE-shaped, so existing behavior against providers/mocks
that ignore stream:true is unaffected) before going through the existing
validation/error-handling.

Verified against DigitalOcean's inference-proxy (kimi-k3): a prompt that
previously hung indefinitely non-streaming now completes in ~9 minutes
streamed (26k+ reasoning tokens).
Persist benchmark lifecycle metadata and artifacts while enabling GPQA and TAU runs through an authenticated API and dashboard.

Co-authored-by: Cursor <cursoragent@cursor.com>
Retain the API/dashboard, custom inference endpoints, TAU simulator configuration, and request auditing while integrating upstream caching, provider controls, host benchmarks, and benchmark updates.

Co-authored-by: Cursor <cursoragent@cursor.com>
@pvennacherla
pvennacherla merged commit 7e1b433 into main Sep 30, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants