Skip to main content
The TraceRoot CLI reads the evaluation data the web app shows — your datasets, the versions you published, and the runs that scored them — without leaving the terminal where the eval ran.
A project API key is already scoped to one project. A browser login identifies you rather than a project, so pass --project <id> (or set TRACEROOT_PROJECT_ID) — traceroot projects list shows the ids.

Find a run

Start at the evaluation, narrow to its runs, then read one. Each step prints the id the next one takes.
--status takes running, completed, completed_with_errors, failed, incomplete, or cancelled. Both lists take --name <substring> filtering, case-insensitively, on the evaluation name. A run row is identity and outcome only — which run it is, what it was testing, and when it started. Scores and counts are aggregates over a run’s results, so they come from reading the run itself.

Read a run

Each score is a mean over the results that carried it, not a headline number for the run:
  • A numeric score’s value is its mean, and CASES is how many results it was taken over.
  • A boolean score’s value is the share that came back true, marked (share true).
  • A categorical score has no mean. It prints categorical rather than a number — a label has no average.
  • Cost and duration are labelled (mean per case). A run’s total cost is a different, and usually much larger, number.

Read a dataset

Datasets read from the dataset down to the cases inside one published version.
datasets versions get prints the version’s cases — case id, input, expected, and the trace a case was captured from where it has one. Versions are immutable, so this is exactly what a run measured: take dataset_version_id off a run and read the cases behind it. A dataset with nothing published yet has no version to read. The CLI says so instead of printing an empty table — publish one from the SDK first, as in Datasets.

Scripting

Every command takes --json and prints the API response verbatim on stdout. Warnings and counts go to stderr, so a redirect stays valid JSON.
Commands exit 0 on success and with a class-specific code otherwise — 2 usage, 3 auth, 4 not found, 5 network — so a script can retry a blip and give up on a missing run without parsing prose.

Two things to know about the output

An absent value prints —, never 0. A run that is still going has no scored, task-error or scorer-error count yet, and the em dash says exactly that; a 0 would claim a measurement nobody made. (none) is a third thing again: a dataset that exists and has no published version. Each read returns one page. There is no cursor flag, so when more matched than came back, the command says so on stderr and names the ceiling. The lists return 50 by default and take --limit up to 200; datasets versions get reads 200 cases by default and takes --limit up to 1000. To pull a whole dataset version into code, use pull_dataset_version() from the SDK rather than the CLI.

Command reference

Next steps

Reading Results

The summary and per-case results evaluate() returns in process.

CLI

Traces, exports, and the rest of what the CLI reads.