Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 28 additions & 2 deletions API.md
Original file line number Diff line number Diff line change
Expand Up @@ -145,6 +145,29 @@ For SWE Atlas runs, `judgeModel` may also be provided:
The judge always uses the server-configured DigitalOcean inference endpoint and
key. `judgeModel` is ignored for non-SWE-Atlas benchmarks.

TAU Airline always uses `openai-gpt-5.4-mini` for its simulated customer through a
DigitalOcean inference endpoint. When `inference.baseUrl` is
`https://inference.do-ai.run/v1` or
`https://inference.do-ai-test.run/v1`, the simulator reuses that endpoint and
`inference.apiKey`. For OpenRouter or a custom candidate endpoint, provide a
separate top-level `simulatorApiKey`; the simulator then uses
`https://inference.do-ai.run/v1`:

```json
{
"benchmark": "tau_bench_verified_airline",
"simulatorApiKey": "<DIGITALOCEAN_SIMULATOR_ACCESS_TOKEN>",
"inference": {
"baseUrl": "https://openrouter.ai/api/v1",
"apiKey": "<OPENROUTER_API_KEY>",
"model": "provider/model"
}
}
```

Both inference secrets are launch-only and are never returned, logged, or
persisted.

Inference constraints:

- `model` accepts letters, digits, `.`, `_`, `:`, `-`, and `/`.
Expand Down Expand Up @@ -595,11 +618,14 @@ Request:
{
"sampleId": "sample-id",
"originalEpoch": 0,
"apiKey": "<INFERENCE_API_KEY>"
"apiKey": "<INFERENCE_API_KEY>",
"simulatorApiKey": "<DIGITALOCEAN_SIMULATOR_ACCESS_TOKEN>"
}
```

The source item must exist in the benchmark-specific report.
`simulatorApiKey` is required only when retrying a TAU Airline run whose
candidate used OpenRouter or a custom endpoint. The source item must exist in
the benchmark-specific report.

Response: `202 Accepted`

Expand Down
8 changes: 1 addition & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -136,12 +136,6 @@ export DO_MODEL_CATALOG_TOKEN='replace-me'
# OpenRouter API key used only to list OpenRouter models.
# Its successful catalog response is cached by the server for two hours.
export OPENROUTER_MODEL_CATALOG_TOKEN='replace-me'
# Optional TAU Airline user-simulator endpoint. When the API key is set, TAU
# uses this independent OpenAI-compatible endpoint for its simulated customer.
export TAU_AIRLINE_USER_SIMULATOR_API_KEY='replace-me'
export TAU_AIRLINE_USER_SIMULATOR_BASE_URL='https://generativelanguage.googleapis.com/v1beta/openai'
export TAU_AIRLINE_USER_SIMULATOR_MODEL='gemini-2.5-flash'

export MYSQL_HOST='replace-me.db.ondigitalocean.com'
export MYSQL_PORT='25060'
export MYSQL_USER='doadmin'
Expand Down Expand Up @@ -220,7 +214,7 @@ curl -X POST http://127.0.0.1:8080/runs \
}'
```

The API and dashboard default GPQA and TAU to three epochs and concurrency three. Deep SWE, SWE-bench Verified, Terminal-Bench, and SWE Atlas default to one epoch and concurrency one because each evaluation provisions disposable sandbox workers. All benchmarks default to reasoning effort `high`, a one-hour full-response timeout per attempt, six retries for retryable inference errors, and retry-on-error behavior. GPQA defaults to temperature `1`; TAU, Deep SWE, SWE-bench Verified, Terminal-Bench, and SWE Atlas default to temperature `0`. Request timeouts (`408`), rate limits (`429`), transport failures, and server errors (`5xx`) are retryable; `maxRetries: 0` disables inference retries. Every value can be overridden. Set `"benchmark": "tau_bench_verified_airline"` to run the 50-task TAU suite, `"benchmark": "deep_swe"` for the 113-task Deep SWE suite, `"benchmark": "swe_bench_verified"` for the 500-task SWE-bench Verified suite, `"benchmark": "terminal_bench"` for the 89-task Terminal-Bench 2.1 suite, or `"benchmark": "swe_atlas_qa"` for the 124-task SWE Atlas QA suite. SWE Atlas runs accept a top-level `"judgeModel"` selected independently from `inference.model`; the judge always uses `https://inference.do-ai.run/v1` and the server-side `SWE_ATLAS_JUDGE_API_KEY`. The dashboard shows this field only for SWE Atlas. SWE-bench uses Harbor's deterministic verifier and does not use a judge model. Each SWE-bench task uses one worker for both candidate work and verification; hidden tests are uploaded only after the candidate patch is captured. The patch and verifier report are saved in sample metadata in the result Parquet and therefore included in the normal Spaces run bundle. Terminal-Bench and SWE Atlas also use one worker per concurrent evaluation, with agent and verifier sharing the worker. SWE Atlas task metadata requests 16 CPUs and 16 GiB RAM, so configure its dedicated size override before increasing concurrency. If `TAU_AIRLINE_USER_SIMULATOR_API_KEY` is configured, its model, base URL, and key are used for TAU's simulated customer while `inference` continues to configure the evaluated agent model; its default model is `gemini-2.5-flash`, and `TAU_AIRLINE_USER_SIMULATOR_MODEL` overrides it. Optional inference fields are `temperature`, `endpointId`, `costTier`, `sort`, `providerOnly`, `allowFallbacks`, `cloudflareVersion`, `costQualityTradeoff`, `pinModel`, and `completionTimeoutMs`. `completionTimeoutMs` limits the complete inference attempt, including connection and streamed-body processing. The existing `timeoutMs` remains the connection/response-header timeout. When OpenRouter is selected, the dashboard exposes provider routing only inside Advanced configuration. Selecting DigitalOcean sends `providerOnly: ["digitalocean"]` and disables provider fallback by default, so OpenRouter must use DigitalOcean or fail the request. Execution can use `start` plus either `end` or `limit`. Set `execution.unordered` to `true` for rolling concurrency without input-order head-of-line blocking; omitted or `false` preserves ordered result emission.
The API and dashboard default GPQA and TAU to three epochs and concurrency three. Deep SWE, SWE-bench Verified, Terminal-Bench, and SWE Atlas default to one epoch and concurrency one because each evaluation provisions disposable sandbox workers. All benchmarks default to reasoning effort `high`, a one-hour full-response timeout per attempt, six retries for retryable inference errors, and retry-on-error behavior. GPQA defaults to temperature `1`; TAU, Deep SWE, SWE-bench Verified, Terminal-Bench, and SWE Atlas default to temperature `0`. Request timeouts (`408`), rate limits (`429`), transport failures, and server errors (`5xx`) are retryable; `maxRetries: 0` disables inference retries. Every value can be overridden. Set `"benchmark": "tau_bench_verified_airline"` to run the 50-task TAU suite, `"benchmark": "deep_swe"` for the 113-task Deep SWE suite, `"benchmark": "swe_bench_verified"` for the 500-task SWE-bench Verified suite, `"benchmark": "terminal_bench"` for the 89-task Terminal-Bench 2.1 suite, or `"benchmark": "swe_atlas_qa"` for the 124-task SWE Atlas QA suite. SWE Atlas runs accept a top-level `"judgeModel"` selected independently from `inference.model`; the judge always uses `https://inference.do-ai.run/v1` and the server-side `SWE_ATLAS_JUDGE_API_KEY`. The dashboard shows this field only for SWE Atlas. SWE-bench uses Harbor's deterministic verifier and does not use a judge model. Each SWE-bench task uses one worker for both candidate work and verification; hidden tests are uploaded only after the candidate patch is captured. The patch and verifier report are saved in sample metadata in the result Parquet and therefore included in the normal Spaces run bundle. Terminal-Bench and SWE Atlas also use one worker per concurrent evaluation, with agent and verifier sharing the worker. SWE Atlas task metadata requests 16 CPUs and 16 GiB RAM, so configure its dedicated size override before increasing concurrency. TAU Airline always uses `openai-gpt-5.4-mini` for its simulated customer through DigitalOcean inference. A production or test DigitalOcean candidate reuses its `inference.baseUrl` and `inference.apiKey`; an OpenRouter or custom candidate requires a separate top-level `simulatorApiKey`, which is used only with `https://inference.do-ai.run/v1`. The dashboard requests this token only when required, and neither secret is persisted. Optional inference fields are `temperature`, `endpointId`, `costTier`, `sort`, `providerOnly`, `allowFallbacks`, `cloudflareVersion`, `costQualityTradeoff`, `pinModel`, and `completionTimeoutMs`. `completionTimeoutMs` limits the complete inference attempt, including connection and streamed-body processing. The existing `timeoutMs` remains the connection/response-header timeout. When OpenRouter is selected, the dashboard exposes provider routing only inside Advanced configuration. Selecting DigitalOcean sends `providerOnly: ["digitalocean"]` and disables provider fallback by default, so OpenRouter must use DigitalOcean or fail the request. Execution can use `start` plus either `end` or `limit`. Set `execution.unordered` to `true` for rolling concurrency without input-order head-of-line blocking; omitted or `false` preserves ordered result emission.

Use `"swe_atlas_qa"` for the 124-task Codebase Q&A track, `"swe_atlas_tw"` for the 90-task Test Writing track, or `"swe_atlas_rf"` for the 65-task Refactoring track. All three use the same independently selected judge model and default to one epoch and concurrency one in managed runs.

Expand Down
6 changes: 1 addition & 5 deletions src/benchmarks/benchmark-config.ts
Original file line number Diff line number Diff line change
Expand Up @@ -16,10 +16,7 @@ import {
ORI_CHANNELS,
ORI_REASONING_EFFORTS,
} from "./agent-cli/schema";
import {
TAU3_BENCH_BANKING_META,
TAU_BENCH_AIRLINE_META,
} from "./benchmark-meta";
import { TAU3_BENCH_BANKING_META } from "./benchmark-meta";
import { DEFAULT_STEP_LIMIT as DEEP_SWE_DEFAULT_STEP_LIMIT } from "./deep-swe/schema";
import { DracoPanelConfigSchema } from "./draco/schemas";
import { SearchLaneConfigSchema } from "./search/core/config";
Expand Down Expand Up @@ -99,7 +96,6 @@ export type MmluProBenchmarkConfig = z.infer<
>;

export const TauBenchOptionsSchema = z.object({
userModel: zDefaultedText(TAU_BENCH_AIRLINE_META.userModel),
userReasoningEffort: z.enum(REASONING_EFFORTS).default("medium"),
});

Expand Down
2 changes: 1 addition & 1 deletion src/benchmarks/benchmark-meta.ts
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ export const TAU_BENCH_AIRLINE_META = {
id: "tau_bench_verified_airline",
defaultEpochs: 3,
temperature: 0,
userModel: "openai/gpt-5.4-mini",
userModel: "openai-gpt-5.4-mini",
} as const satisfies BenchmarkMeta;

export const TAU3_BENCH_BANKING_META = {
Expand Down
19 changes: 4 additions & 15 deletions src/benchmarks/tau-bench-airline/airline.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -26,8 +26,8 @@ import { benchmarkIds, getBenchmark } from "../registry";
import { compareActionWithToolCall } from "./action-match";
import {
airlineRecordToSample,
resolveAirlineUserModel,
TAU_BENCH_AIRLINE_ID,
TAU_BENCH_AIRLINE_USER_SIMULATOR_MODEL,
} from "./benchmark";
import {
ensureAirlineData,
Expand Down Expand Up @@ -120,24 +120,13 @@ describe("tau_bench_verified_airline registry", () => {
expect(b?.id).toBe("tau_bench_verified_airline");
expect(b?.temperature).toBe(0);
expect(b?.defaultEpochs).toBe(3);
expect(b?.userModel).toBe("openai/gpt-5.4-mini");
expect(b?.userModel).toBe("openai-gpt-5.4-mini");
});
it("appears in benchmarkIds()", () => {
expect(benchmarkIds()).toContain("tau_bench_verified_airline");
});
it("uses the DigitalOcean user-model slug only for DigitalOcean inference", () => {
expect(
resolveAirlineUserModel(
"openai/gpt-5.4-mini",
"https://inference.do-ai.run/v1"
)
).toBe("openai-gpt-5.4-mini");
expect(
resolveAirlineUserModel(
"openai/gpt-5.4-mini",
"https://openrouter.ai/api/v1"
)
).toBe("openai/gpt-5.4-mini");
it("fixes the user simulator to the DigitalOcean model slug", () => {
expect(TAU_BENCH_AIRLINE_USER_SIMULATOR_MODEL).toBe("openai-gpt-5.4-mini");
});
});
describe("tau_bench_verified_airline config", () => {
Expand Down
38 changes: 14 additions & 24 deletions src/benchmarks/tau-bench-airline/benchmark.ts
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@ import { Solver } from "../../harness/solver";
import { Either } from "../../internal/either";
import { definedValues } from "../../internal/guards";
import { parseSchema } from "../../internal/zod";
import { isDigitalOceanInferenceBaseUrl } from "../../providers/digitalocean-inference";
import { makeOpenRouterModelLayer } from "../../providers/openrouter-model";
import {
makeResponsesModelLayer,
Expand All @@ -40,25 +41,8 @@ export const TAU_BENCH_AIRLINE_TEMPERATURE = TAU_BENCH_AIRLINE_META.temperature;

export const TAU_BENCH_AIRLINE_ID = TAU_BENCH_AIRLINE_META.id;

const DIGITALOCEAN_INFERENCE_BASE_URLS = new Set([
"https://inference.do-ai.run/v1",
"https://inference.do-ai-test.run/v1",
]);

export const DIGITALOCEAN_MODEL_SLUGS: Readonly<Record<string, string>> = {
"openai/gpt-5.4-mini": "openai-gpt-5.4-mini",
};

export function resolveAirlineUserModel(
model: string,
baseUrl: string | undefined
): string {
const normalizedBaseUrl = baseUrl?.replace(/\/+$/, "");
return normalizedBaseUrl !== undefined &&
DIGITALOCEAN_INFERENCE_BASE_URLS.has(normalizedBaseUrl)
? (DIGITALOCEAN_MODEL_SLUGS[model] ?? model)
: model;
}
export const TAU_BENCH_AIRLINE_USER_SIMULATOR_MODEL =
TAU_BENCH_AIRLINE_META.userModel;

export function airlineRecordToSample(
record: Readonly<Record<string, unknown>>,
Expand Down Expand Up @@ -123,11 +107,17 @@ function makeAirlineLayer(
}
const userSimulator = input.userSimulator;
const userSimulatorBaseUrl = userSimulator?.baseUrl ?? input.baseUrl;
const defaultUserModel = resolveAirlineUserModel(
benchmarkConfig.userModel,
userSimulatorBaseUrl
);
const userSimulatorModel = userSimulator?.model ?? defaultUserModel;
if (
userSimulatorBaseUrl === undefined ||
!isDigitalOceanInferenceBaseUrl(userSimulatorBaseUrl)
) {
return layerFail(
new Error(
"TAU Airline user simulator requires a DigitalOcean inference endpoint and access token"
)
);
}
const userSimulatorModel = TAU_BENCH_AIRLINE_USER_SIMULATOR_MODEL;
const solverOpts: SolverOpts = definedValues({
endpointId: benchmarkConfig.endpointId,
userModelConfig: definedValues({
Expand Down
2 changes: 1 addition & 1 deletion src/benchmarks/tau-bench-airline/user-simulator.ts
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ import { retrySalted, withRetryAttemptLogging } from "../../runtime/retry";
import type { UserModelConfig } from "./types";
import { USER_SIM_GUIDELINES } from "./user-sim-guidelines";

const USER_FALLBACK_MODEL = "openai/gpt-5.4-mini";
const USER_FALLBACK_MODEL = "openai-gpt-5.4-mini";

class UserSimError extends TaggedError("UserSimError")<{
readonly message: string;
Expand Down
1 change: 0 additions & 1 deletion src/benchmarks/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,6 @@ export interface BenchmarkRunInput<
readonly userSimulator?: {
readonly apiKey: string;
readonly baseUrl: string;
readonly model: string;
};
readonly benchmarkConfig: Config;
readonly sessionId: string;
Expand Down
37 changes: 29 additions & 8 deletions src/cli/index.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -368,16 +368,37 @@ describe("bench-harness CLI", () => {
retrievalConfig: "bm25_grep",
});
});
it("uses server-provided Gemini settings for the TAU user simulator", () => {
it("reuses DigitalOcean candidate credentials for the TAU simulator", () => {
for (const baseUrl of [
"https://inference.do-ai.run/v1",
"https://inference.do-ai-test.run/v1/",
]) {
expect(
tauAirlineUserSimulatorFromEnv({}, "candidate-key", baseUrl)
).toEqual({
apiKey: "candidate-key",
baseUrl: baseUrl.replace(/\/+$/u, ""),
});
}
});

it("requires a separate DigitalOcean token for non-DO candidates", () => {
expect(
tauAirlineUserSimulatorFromEnv({
TAU_AIRLINE_USER_SIMULATOR_API_KEY: "gemini-key",
})
tauAirlineUserSimulatorFromEnv(
{ TAU_AIRLINE_USER_SIMULATOR_API_KEY: "simulator-key" },
"candidate-key",
"https://openrouter.ai/api/v1"
)
).toEqual({
apiKey: "gemini-key",
baseUrl: "https://generativelanguage.googleapis.com/v1beta/openai",
model: "gemini-2.5-flash",
apiKey: "simulator-key",
baseUrl: "https://inference.do-ai.run/v1",
});
expect(tauAirlineUserSimulatorFromEnv({})).toBeUndefined();
expect(() =>
tauAirlineUserSimulatorFromEnv(
{},
"candidate-key",
"https://inference.example.com/v1"
)
).toThrow("TAU_AIRLINE_USER_SIMULATOR_API_KEY");
});
});
50 changes: 32 additions & 18 deletions src/cli/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,10 @@ import { runHarnessPromise } from "../internal/effect-logger";
import { Either } from "../internal/either";
import { definedValues, isMember } from "../internal/guards";
import { parseSchema } from "../internal/zod";
import {
DEFAULT_DIGITALOCEAN_INFERENCE_BASE_URL,
isDigitalOceanInferenceBaseUrl,
} from "../providers/digitalocean-inference";
import { makeLocalResultStore } from "../results/result-store";
import { datasetSizeById, runBenchmarkById } from "../runner/run-by-id";

Expand Down Expand Up @@ -176,25 +180,31 @@ function resolveApiKey(): string {
}

export function tauAirlineUserSimulatorFromEnv(
env: NodeJS.ProcessEnv = process.env
):
| {
readonly apiKey: string;
readonly baseUrl: string;
readonly model: string;
}
| undefined {
env: NodeJS.ProcessEnv,
inferenceApiKey: string,
inferenceBaseUrl: string | undefined
): {
readonly apiKey: string;
readonly baseUrl: string;
} {
if (
inferenceBaseUrl !== undefined &&
isDigitalOceanInferenceBaseUrl(inferenceBaseUrl)
) {
return {
apiKey: inferenceApiKey,
baseUrl: inferenceBaseUrl.replace(/\/+$/u, ""),
};
}
const apiKey = env["TAU_AIRLINE_USER_SIMULATOR_API_KEY"]?.trim();
if (!apiKey) {
return undefined;
throw new Error(
"Set TAU_AIRLINE_USER_SIMULATOR_API_KEY when the TAU candidate uses OpenRouter or a custom inference endpoint."
);
}
return {
apiKey,
baseUrl:
env["TAU_AIRLINE_USER_SIMULATOR_BASE_URL"]?.trim() ||
"https://generativelanguage.googleapis.com/v1beta/openai",
model:
env["TAU_AIRLINE_USER_SIMULATOR_MODEL"]?.trim() || "gemini-2.5-flash",
baseUrl: DEFAULT_DIGITALOCEAN_INFERENCE_BASE_URL,
};
}

Expand All @@ -207,13 +217,17 @@ function main(): Promise<void> {
);
}
const apiKey = resolveApiKey();
const tauAirlineUserSimulator =
args.benchmark === "tau_bench_verified_airline"
? tauAirlineUserSimulatorFromEnv()
: undefined;
const baseUrl = getOrNull(
runSync(string("OPENROUTER_BASE_URL").pipe(option))
);
const tauAirlineUserSimulator =
args.benchmark === "tau_bench_verified_airline"
? tauAirlineUserSimulatorFromEnv(
process.env,
apiKey,
baseUrl ?? undefined
)
: undefined;
const epochs = args.epochs ?? benchmark.defaultEpochs;
const range = resolveRange(args);
const sessionId = resolveSessionId();
Expand Down
18 changes: 18 additions & 0 deletions src/providers/digitalocean-inference.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
export const DIGITALOCEAN_INFERENCE_BASE_URLS = [
"https://inference.do-ai.run/v1",
"https://inference.do-ai-test.run/v1",
] as const;

export const DEFAULT_DIGITALOCEAN_INFERENCE_BASE_URL =
DIGITALOCEAN_INFERENCE_BASE_URLS[0];

function withoutTrailingSlashes(value: string): string {
return value.replace(/\/+$/u, "");
}

export function isDigitalOceanInferenceBaseUrl(value: string): boolean {
const normalized = withoutTrailingSlashes(value);
return DIGITALOCEAN_INFERENCE_BASE_URLS.some(
(baseUrl) => baseUrl === normalized
);
}
1 change: 0 additions & 1 deletion src/runner/run-by-id.ts
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,6 @@ export interface RunBenchmarkInput {
readonly userSimulator?: {
readonly apiKey: string;
readonly baseUrl: string;
readonly model: string;
};
readonly benchmarkConfig: BenchmarkRunConfig;
readonly epochs: number;
Expand Down
Loading
Loading