Evals @ Prime Intellect
-
Prime Intellect
- Germany
-
16:12
(UTC +02:00) - https://florianbrand.com/
- @xeophon
Highlights
- Pro
Popular repositories Loading
-
-
yet-another-applied-llm-benchmark
yet-another-applied-llm-benchmark PublicForked from carlini/yet-another-applied-llm-benchmark
A benchmark to evaluate language models on questions I've previously asked them to solve.
Python 2
-
AutomationBench-PI
AutomationBench-PI PublicForked from zapier/AutomationBench
A benchmark for evaluating AI agents on realistic business workflows
-
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.





