Skip to content
Discussion options

You must be logged in to vote

Hello @Parkash007-max,

Apologies for the delay in our response.

(1) The error bars in the figure represent the macro-averaged standard error across 10 runs per task. Although we did not vary random seeds, the system is non-deterministic even under fixed seed and temperature settings. This is due to variability in the input derived from live observability data, such as timestamps, which introduces natural variance across runs. The standard error reflects this inherent variability in the model's output behavior.

(2) In the main text, the combined average of the agent's performance on the same 21 scenarios with and without access to traces has been reported. Hence the number 42 (21x2). In th…

Replies: 1 comment

Comment options

You must be logged in to vote
0 replies
Answer selected by Red-GV
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
2 participants
Converted from issue

This discussion was converted from issue #20 on January 08, 2026 14:20.