# evalshift-action (GitHub Action) — complete reference for AI tools Canonical hosted copy: https://www.evalshift.dev/ci-llms-full.txt Repo/action ref: evalshift/evalshift-action | version: 0.6.0 | license: MIT Kind: composite GitHub Action (not JavaScript, not Docker) Marketplace name: "EvalShift" | branding: bar-chart-2 / purple Runtime helper: scripts/evalshift_action.py (stdlib only, no third-party deps) Purpose: check the org's plan covers this job before spending credits, run an EvalShift golden suite inside CI, push the completed run to hosted EvalShift, ask the hosted API for a compatible baseline run on the base branch, fetch the hosted diff, ask the hosted governed gate to judge the run against the migration policy it was pushed with (the `migration_policy` block of `evalshift.yaml`, snapshotted into the run), write action outputs, keep exactly one PR comment up to date, set the `evalshift/regression` commit status, and exit non-zero when the selected `fail-on` mode says to. The action is a thin CI wrapper: all evaluation, statistics, and reporting happen in the EvalShift CLI it installs; all diffing happens server-side in hosted EvalShift. Related packages (separate repos, separate docs): - evalshift (PyPI, the CLI this action installs and shells out to) — https://www.evalshift.dev/cli-llms-full.txt - evalshift-sdk (PyPI, in-process capture SDK) — https://www.evalshift.dev/sdk-llms-full.txt Minimal usage: ```yaml permissions: contents: read pull-requests: write issues: write statuses: write jobs: evalshift: runs-on: ubuntu-latest env: EVALSHIFT_NONINTERACTIVE: "1" ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} steps: - uses: actions/checkout@v7 - uses: evalshift/evalshift-action@v0 with: token: ${{ secrets.EVALSHIFT_TOKEN }} fail-on: policy # the default ``` Prerequisites in the repository, all of them hard requirements: 1. `evalshift.yaml` committed (or `config:` pointed at wherever it lives). 2. A golden JSONL suite committed, selected with `suite-name:` (a `suites:` key) or `suite:` (a path). 3. Repository (or environment) secret `EVALSHIFT_TOKEN` — an `es_...` service-account key, scoped to `run:create` + `run:read` + `policy:read`. See "Token hygiene" below. Without `policy:read` the default `fail-on: policy` gate falls back to regression mode on every run (a stderr warning, not a failure — easy to miss). 4. A model provider key in the job env matching the models named in `evalshift.yaml`. Every run makes real model calls and spends real credits. Cost per run ≈ suite size × 2 models (source + target) × prompts, minus CLI cache hits (the runner starts with a cold cache, so in practice assume no cache reuse across CI runs). ## Execution model Composite `runs.steps`, in order: 1. `actions/setup-python@v7` with `python-version` (default `3.12`). 2. `python -m pip install --upgrade pip` then `python -m pip install "evalshift=="`. No pip caching; expect ~20-60s of install per run. 3. `python "$GITHUB_ACTION_PATH/scripts/evalshift_action.py"` with `id: run`, all inputs passed as `INPUT_*` env vars. (Composite actions do NOT auto-populate `INPUT_*` for `run:` steps, so action.yml maps every input explicitly — an input added to `inputs:` without a matching `env:` entry is invisible to the helper.) The helper's process CWD is the workspace root (`Path.cwd()`), so `config` / `suite` paths and the `.evalshift/` output directory are all relative to the repository root by default. The helper never streams: `run_command` uses `subprocess.run(capture_output=True)` and prints stdout/stderr only after each CLI command finishes. A long `evalshift all` looks silent in the log until it completes. ## Inputs (action.yml) | Input | Required | Default | Meaning | |---|---|---|---| | token | yes | — | Hosted EvalShift API token (`es_...`). Masked via `::add-mask::`, and redacted from CLI output. Passed to the CLI as env `EVALSHIFT_TOKEN`, never in argv. | | host | no | https://api.evalshift.dev | Hosted API base URL. Trailing `/` stripped. Passed as env `EVALSHIFT_HOST`. Set only for self-hosted/staging. | | config | no | evalshift.yaml | Path to config, relative to workspace root. Paths *inside* the config resolve relative to the config file's own directory (CLI behavior), so a config in a subdirectory works unchanged. | | suite | no | golden.jsonl | Path to golden JSONL suite, relative to workspace root. Selects the FILE only. Mutually exclusive with suite-name. | | suite-name | no | (empty) | `suites:` key in evalshift.yaml. Carries that suite's own evaluators block. Needs evalshift-version >= 0.14.0. Mutually exclusive with suite. | | evalshift-version | no | 1.2.0 | Exact CLI version installed from PyPI. Pinned for reproducibility. | | python-version | no | 3.12 | Python for the CLI. Must satisfy the CLI's `requires-python` (>=3.11 for 1.2.0); lowering it below that breaks install. | | fail-on | no | policy | `policy` \| `never` \| `regression` \| `any-slice-regression`. Any other value → hard error before any work. DEFAULT CHANGED: was `regression`; `policy` gates on the server's migration-policy verdict, not on the diff. | | require-policy | no | false | Whether `inconclusive` + `policy_source: "none"` (a run pushed with no `migration_policy`) fails the job. Default false: reported as ungated — a `::warning::` annotation on stdout and a commit status saying the gate is off — and the job still passes. Read under `fail-on: policy` only; it never overrides a verdict the policy did reach, and never the unavailable-check fallback. | | branch | no | "" (auto) | Candidate branch recorded on the hosted run. Auto-detected; override only for non-standard naming. | | base-branch | no | "" (auto) | Branch searched for a baseline run. Auto-detected. If it resolves empty, no baseline lookup happens at all; that passes the diff-based `fail-on` modes (`regression`, `any-slice-regression`) but does not affect `fail-on: policy` (the default), which gates on the run's own policy check regardless of baseline. | | create-project | no | true | When false, appends `--no-create-project` to `evalshift push`, making a missing hosted project a hard failure. | | comment | no | true | Whether to upsert the PR comment. Commit status is set regardless. | | github-token | no | `${{ github.token }}` | Token for the PR comment and commit status. Also masked. If empty, both comment and status are skipped entirely. | | repo-private | no | `${{ github.event.repository.private }}` | Repository visibility, asserted to the CI preflight (never verified server-side). Drives the `private_repo_ci` entitlement check. | Boolean inputs are truthy for `1`, `true`, `yes`, `on` (case-insensitive); everything else is false. All inputs are `.strip()`ed. ## Outputs | Output | Value | |---|---| | run_url | Hosted run URL, parsed as the last http(s) line of `evalshift push` stdout. Empty is impossible — a push that prints no URL is a hard error. | | diff_url | Hosted `web_diff_url` for the baseline comparison; empty string when no compatible baseline. | | run_id | Hosted run id — the server-minted UUID, read out of the `/app/{org}/{project}/runs/` path of `run_url`. NOT the local run directory name (`r_20260723_mysuite_ab12cd`), which addresses nothing beyond the runner's own disk. | | regression_count | `aggregate_delta.regressions` from the hosted diff; `0` when no baseline. | | conclusion | `success` or `failure`, after applying `fail-on`. It is a GitHub commit-status state, so both a `conditional_pass` and an undecided policy (`inconclusive` or an unknown status) read `success` here; the PR comment and the status description carry the verdict. | Written by appending `key=value` lines to `$GITHUB_OUTPUT`. Values are single-line by construction; there is no heredoc delimiter handling, so a hypothetical multi-line value would corrupt the output file. ## Commands the action shells out to ``` evalshift all --yes --config evalshift push --config [--no-create-project] # is `--suite-name ` when the suite-name input is set, else `--suite `. ``` `all` is deliberate, not stale. The CLI renamed that command to `evalshift compare` and kept `all` registered permanently as a hidden alias bound to the same function — no deprecation, no removal date. The action types `all` because it is the one spelling EVERY installable CLI version answers to, the pre-rename releases included (the pin defaults to one, and users may pin older still). The only difference at runtime is a one-line rename notice the CLI prints on stderr; stdout, exit codes and artefacts are identical. Do not rewrite this invocation to `compare` unless the minimum supported pin moves past the release that introduced it. Human docs say `compare` for what a developer types locally. Exactly one suite-selection flag is passed, and both commands get the same one, so the CLI's ambiguous-suite error — and the framed panel of per-suite commands it prints when a config wires several `suites:` entries — cannot appear in CI. The default, when neither input is set, is the literal path `golden.jsonl`, not "whichever suite is wired". NAME vs PATH is not cosmetic. `--suite ` selects a file; `--suite-name ` also resolves that suite's own `evaluators:` override out of `suites:` (CLI: `EvalShiftConfig.evaluators_for`, which maps a `None` name to the top-level block). Select a wired suite by path and it is silently scored with the top-level evaluators — for a tool-calling suite under a `semantic` + `llm_judge` top level that scores zero rows and the run dies at `analyze` with `scores.jsonl is empty`. So: a suite with an entry under `suites:` is selected with `suite-name`; `suite` is for a file that is not wired into the config. Setting both inputs is refused (`ActionError`, "mutually exclusive") rather than resolved by precedence, and `suite-name` on a pin below 0.14.0 — the release that added `--suite-name` — is refused with the pin named. Both commands inherit the job environment plus `EVALSHIFT_HOST`, `EVALSHIFT_TOKEN` and `COLUMNS=512`. A non-zero exit from either raises `ActionError` and fails the step (message: `command failed (): `). `COLUMNS` is set because the CLI prints through rich, which folds output at the console width (80 when stdout is a pipe, which it always is here) and would otherwise split the hosted run URL across two lines. The `` argument to `push` is the LOCAL run id: the directory under `/.evalshift/runs` with the newest mtime, never parsed from CLI output. Consequences: a pre-existing `.evalshift/runs` in the checkout is harmless (the fresh run is newest), but any step that touches an older run directory between `all` and the run-id read can select the wrong run. Missing directory → `run directory ... does not exist`; empty directory → `no local EvalShift runs found in ...`. `run_url` extraction scans `push` stdout bottom-up for the first non-empty line starting with `http://` or `https://`. A CLI release that stops printing a bare URL as the last line breaks this (guarded by the `cli-contract` CI job, which only checks flags, not output shape). The HOSTED run id — the `run_id` output, and the id in every `/runs/{id}` call below — is the path segment immediately after `/runs/` in that `run_url`, because the server mints it and the URL is the only place `push` reports it. The server builds exactly one shape, `{web_app_url}/app/{org}/{project}/runs/{id}`, and the whole of it is anchored — not just the `/runs/` marker, and not simply taken off the end. A `.../projects/` URL is the right shape and the wrong thing; `.../runs//diff/` ends on the other side of the comparison; and an org slugged `runs` in front of a UUID-shaped project slug puts a decoy `/runs/<36 chars>` to the LEFT of the real one. Anchored but not rooted, so a `web_app_url` carrying a base path still matches. It must match the canonical UUID shape or the step fails with `could not read a server run id out of the hosted run URL: ...`; guessing would buy a 404 on `policy-check`, which the gate reads as "no stored policy decision" and degrades past. The local `r_…` id is never sent to the API. ## Plan preflight (runs BEFORE the CLI) Purpose: refuse a job the org's plan does not cover before any model credits are spent. Sequence, before `evalshift all`: 1. `project_ref_from_config(Path(config))` — regex `^project:\s*["']?([a-z0-9-]+)/([a-z0-9-]+)` against the config file, one line at a time. No YAML parser (helper is dependency-free). No match (or unreadable file) → preflight skipped entirely, run proceeds. 2. ONE call: `POST {host}/runs/preflight` with body `{"project_slug": "/", "repo_private": , "parallelism": 1}`. Addressed by the same slug `POST /runs` takes; authorized by `run:create` only (the permission the push needs anyway), so any key that can push can ask, project-pinned keys included. No project lookup precedes it. `PREFLIGHT_PARALLELISM = 1` is a constant, not an input: one EvalShift run per job, and the server counts in-flight runs itself. What the server checks, in order: records `repo_private` (sticky); `private_repo_ci`; declared `parallelism` vs `max_parallelism`; `runs_per_month` quota (same meter as `POST /runs`); subscription status and seat overage. In-flight uploads are NOT checked (upload-time only). Outcomes (`run_preflight(hosted, *, project_ref, repo_private, create_project)`): - `2xx` → continue. - `402` → `PreflightDenied(message, details)`, job fails immediately with exit 1, `evalshift` never runs. `details` is the server's envelope: `feature`, `tier`, `limit`, `used`, `status`, `resets_at`, `upgrade_url`. The action renders it and decides nothing itself. - `401` → STOP: `hosted EvalShift at rejected the token (HTTP 401): ...` plus key advice. Exit 1, `evalshift` never runs. - `403` → STOP: `hosted EvalShift refused the plan preflight (HTTP 403): ...` plus `The EvalShift token is missing the 'run:create' permission. ...`. Exit 1. - `404` + `create-project: false` → STOP: `hosted EvalShift has no project '/' this token can reach (HTTP 404), and create-project: false forbids the push to create it. ...`. Exit 1. (404 is existence-hiding: unknown org, unknown project, non-member, or a key pinned to another project all look the same.) - `404` + `create-project: true` → stdout `::notice title=EvalShift preflight::project '/' is not on hosted EvalShift yet, or this token cannot see it (HTTP 404); the first push creates it, ...`, run proceeds. - `405` (server predates the route) → stderr `warning: plan preflight skipped: hosted EvalShift at predates POST /runs/preflight (HTTP 405); ...`, run proceeds. NO fallback to any older preflight flow. - Anything else (5xx, 422, other 4xx, timeout, DNS/URLError, malformed body) → `warning: plan preflight skipped: ` on stderr, run proceeds. Also `http.client.HTTPException` (BadStatusLine, IncompleteRead). Fail-closed on billing and on answers the push would repeat; fail-open on infrastructure. Every STOP is reported like a 402 denial: stdout `::error title=EvalShift::` (one line, the fix after an escaped `%0A`), and a `## EvalShift did not run` block with the message appended to `$GITHUB_STEP_SUMMARY`. No PR comment. Outputs are the same as on a 402 denial (below). Nothing is printed to stderr for a STOP. Denial rendering, all three from the same markdown body (`build_preflight_body`): - stdout: `::error title=EvalShift::%0AUpgrade: ` (workflow commands are single-line; in the message `%`/`\r`/`\n` escape to `%25`/`%0D`/`%0A`, `%` first; the `title=` property additionally escapes `:`/`,` to `%3A`/`%2C`). - `$GITHUB_STEP_SUMMARY`: the body, appended. - PR comment: the same body via `upsert_pr_comment`, carrying `COMMENT_MARKER`, so it replaces the regular EvalShift comment rather than stacking. Requires `github-token` + `comment: true`. Outputs on denial: `run_url=""`, `diff_url=""`, `run_id=""`, `regression_count=0`, `conclusion=failure`. `repo_private` is client-asserted; the server cannot verify it and records the first `true` permanently (set-once), so a later `false` from the same project changes nothing. ## Hosted API contract used Auth on all calls: `Authorization: Bearer `, `Accept: application/json`, 30s timeout, stdlib `urllib`. The preflight call above uses the same client and headers. 1. `GET {host}/runs/{run_id}/baseline-compatible?branch={base_branch}` Response (object required, else hard error): ```json { "baseline_run": {"id": "..."} | null, "compatibility": "direct", "api_diff_url": "/runs//diff/", "web_diff_url": "https://app.evalshift.dev/..." } ``` Skipped entirely when `base_branch` is empty. 2. `GET api_diff_url` (absolute URLs used as-is; relative joined onto `host`). Only called when `api_diff_url` is truthy. Response fields the action reads: ```json { "aggregate_delta": {"regressions": , "pass_rate_delta": }, "per_slice_deltas": [{"slice": "", "pass_rate_delta": }] } ``` Any other key is ignored. Non-object response → hard error. 3. `GET {host}/runs/{run_id}/policy-check` — the governed gate. Called ONLY when `fail-on: policy`, always for the run just pushed, independent of whether a baseline exists. Needs `policy:read`; a 403 is NOT fatal — it falls back to regression gating with `warning: hosted policy check failed: hosted EvalShift refused the request (HTTP 403): Permission denied: policy:read ...` on stderr. ```json { "run_id": "...", "status": "pass" | "conditional_pass" | "fail" | "inconclusive", "verdict": "" | null, "reason": "", "policy_source": "run_policy" | "project_policy" | "none", "policy": { ... }, "budgets": [{"name": "...", "observed": , "allowed": , "passed": , "scope": "...", "ci_low": , "ci_high": , "conclusive": }], "blocking_regressions": [{"prompt_id": "...", "evaluator_name": "...", "slice_name": "...", "severity": "...", "delta_avg_score": , "effect_size": }] } ``` `status` is a CLOSED set of four, and the action recognises all four: - `pass` — every budget within policy. - `conditional_pass` — every budget held and nothing critical/high regressed, but medium/low regressions and/or comparisons that scored zero pairs. A PASSING state with caveats; it deliberately does not fail the gate. Its `reason` contains "not a gate failure" and ends "Review before merging." - `fail` — a budget busted, or a blocking critical/high regression. - `inconclusive` — six distinct causes, distinguishable ONLY by `reason`: (1) the run carried no policy, which returns `policy: null`, `budgets: []` and `policy_source: "none"` — the one cause the action words itself (see below); (2) no policy metrics recorded for the run; (3) nothing measured — no blocking evaluator scored a record, so the quality budgets are clean by absence; (4) nothing comparable — every comparison scored severity `insufficient`, which also returns `budgets: []`; (5) a budget breached on too small a sample to confirm it; (6) a policy-declared slice went unmeasured. The action prints `reason` verbatim and never paraphrases it, or the six collapse into one on the surface readers see. The list has grown before and may again; nothing in the action enumerates it. A string outside those four means a newer server; the action treats it exactly like `inconclusive` (undecided, never a pass) so a pinned version keeps working. `budgets[]` may exceed six entries and `scope` may be a slice name rather than `overall`; it is `[]` for the all-`insufficient` case. Never raises out of `fetch_policy_check` — HTTPError (any code), URLError, a wrapped 403 `ActionError`, a malformed body, or a response whose `status` is missing/blank all return `(None, reason)` and print `warning: ; falling back to fail-on: regression` to stderr. `target_url` for the commit status = `web_diff_url` when present, else `run_url`. ## Gating algorithm (exact) ``` # Diff facts, computed in every mode (the comment renders them even when they do not gate). diff is None -> regression_count=0, slice_regressions=[] regression_count = int(aggregate_delta.regressions or 0) slice_regressions = [s for s in per_slice_deltas if float(s.pass_rate_delta) < 0] sorted ascending by pass_rate_delta # most negative first top_slice_regressions = slice_regressions[:5] # comment display only fail_on == "never" -> should_fail = False fail_on == "regression" -> should_fail = regression_count > 0 fail_on == "any-slice-regression" -> should_fail = len(slice_regressions) > 0 fail_on == "policy": # the default policy-check unavailable -> should_fail = regression_count > 0 # fallback, announced status == "fail" -> should_fail = True status == "pass" -> should_fail = False status == "conditional_pass" -> should_fail = False, reported as a PASS with caveats status == "inconclusive" and policy_source == "none" -> should_fail = require_policy # the run carried no policy ::warning:: annotation on STDOUT, either way status == "inconclusive" -> should_fail = False, reported as undecided any other status -> should_fail = False, reported as unrecognized/undecided conclusion = "failure" if should_fail else "success" ``` Non-numeric / missing deltas coerce to `0.0`, so a malformed slice entry is treated as "flat", never as a regression. `any-slice-regression` is strictly not a superset of `regression`: a run with `regressions > 0` but no negative slice delta fails under `regression` and passes under `any-slice-regression`. `policy` is not a stricter or looser version of `regression` either — it is the server's verdict on the migration policy the run was pushed with, and can disagree in both directions: a diff with regressions that stay inside every budget passes, and a diff with `regressions == 0` that busts a cost/latency/per-slice budget fails. Undecided is never a pass and never a failure: `should_fail=False`, `conclusion=success` (the only GitHub states are success/failure), and `GatingResult.summary` — surfaced in the commit status description and the PR comment — says the gate could not decide, naming the unrecognized status when there is one. An unavailable policy check gates on the diff for that run and is announced in the job log, the status description and the comment; it never silently goes green. `policy_source: "none"` + `inconclusive` is the absence of a gate, not a verdict, and is the only case the action words for itself. `GatingResult.policy_ungated` is True for it, and: - `summary` (commit status + job log + PR comment) reads `the gate is off — no migration policy was pushed with this run; add migration_policy to evalshift.yaml`. The server's `reason` is deliberately NOT appended here (it repeats the same instruction and the status description is cut at 140 chars); it still renders as `**Why:**` in the comment. - a workflow annotation is printed on STDOUT — GitHub scrapes annotations from stdout only, unlike the fallback warning, which goes to stderr: `::warning title=EvalShift policy gate::no migration policy was pushed with this run, so nothing gates this PR; add migration_policy to evalshift.yaml and push again` - the PR comment renders a "Nothing gates this PR" blockquote instead of the "could not decide" one. - with `require-policy: true`, `should_fail=True`, `conclusion=failure`, and `summary` reads `no migration policy was pushed with this run and require-policy is set; add migration_policy to evalshift.yaml`. The annotation is printed either way. `conditional_pass` is NOT undecided. `GatingResult.policy_decided` is True for it (the set is `POLICY_DECIDED_STATUSES = {"pass", "conditional_pass", "fail"}`) and `GatingResult.policy_caveated` is True only for it (`POLICY_CAVEATED_STATUSES = {"conditional_pass"}`). `summary` reads `the gate passed, with caveats: `, and the PR comment renders a "passed, with caveats" blockquote instead of the "could not decide" one. `GatingResult` fields: `conclusion`, `should_fail`, `regression_count`, `top_slice_regressions`, `mode`, `summary`, `policy_status`, `policy_verdict`, `policy_reason`, `policy_source`, `budgets`, `blocking_regressions`, `policy_unavailable_reason`, plus the derived `policy_decided` (`policy_status in {"pass","conditional_pass","fail"}`) and `policy_caveated` (`policy_status in {"conditional_pass"}`). Process exit code: `1` when `should_fail`, `1` on any `ActionError` (printed as `error: ` to stderr), else `0`. ## GitHub context detection ``` event_name = GITHUB_EVENT_NAME repository = GITHUB_REPOSITORY event = JSON at GITHUB_EVENT_PATH (unreadable/invalid -> {}) is_pull_request = event_name.startswith("pull_request") # includes pull_request_target pull_number = event.number if int else None sha = event.pull_request.head.sha or GITHUB_SHA branch = input.branch or event.pull_request.head.ref or GITHUB_HEAD_REF or GITHUB_REF_NAME base_branch = input.base-branch or event.pull_request.base.ref or GITHUB_BASE_REF or GITHUB_REF_NAME ``` On a `push` event, `base_branch` falls back to the pushed branch itself — a push to `main` therefore diffs against the previous `main` run, which is the intended "track the trunk" behavior, not a bug. ## PR comment Marker: `` (first line of the body). Upsert logic: - Skipped unless `comment: true`, `is_pull_request`, and `pull_number is not None`. - `GET /repos/{repo}/issues/{pull_number}/comments`, then the FIRST comment whose author `user.type == "Bot"` AND whose body contains the marker is PATCHed; otherwise a new comment is POSTed. A human-authored comment containing the marker is deliberately never edited. - Pagination is not handled: only the first page of comments is inspected. On a PR with enough comments to push the EvalShift comment off page 1, a duplicate is created. - The marker is a constant, so two invocations of this action on the same PR fight over one comment. Run at most one invocation with `comment: true` per PR. - HTTP 403/404 → `warning: could not upsert PR comment: HTTP ` on stderr and continue. Any other HTTP error propagates and fails the step. Body shape (no baseline): ``` ## EvalShift regression check **Conclusion:** `success` **Hosted run:** [open run]() **Regressions:** 0 No compatible baseline run was found on the base branch. ``` Body shape (with baseline) replaces the trailing line with a `**Diff:**` link (when `web_diff_url` exists), a `**Pass-rate movement:**` line, and a two-column table of up to five regressed slices, or the single row `| No regressed slices | 0 pts |`. Under `fail-on: policy` only, two extra sections sit between the header lines and the diff sections (they render with or without a baseline): - `**Policy decision:** \`\` (from \`\`)` + `**Why:** `, where `` is the server's string verbatim (whitespace-collapsed only). On `conditional_pass`, a blockquote states the gate "passed, with caveats" and points at the reason. On `inconclusive` or an unrecognized status, a blockquote states that the gate could not decide and that this is not a pass; on `inconclusive` with `policy_source: "none"` it states instead that nothing gates this PR and asks for `migration_policy` in `evalshift.yaml`. When the check was unavailable, the whole pair is replaced by a blockquote naming the reason and the `fail-on: regression` fallback. - `### Policy budgets` — `| Budget | Scope | Observed | Allowed | Result |`, failing rows first, capped at `MAX_BUDGET_ROWS = 12`, then ` more budgets not shown — see the hosted run.`, or ` more budgets not shown, of them failing — see the hosted run.` when the hidden rows include failures (per-slice budgets can produce more failures than fit). `Scope` is `overall` or the slice name. Result is `pass`/`fail`/`unknown`, with ` (not confident)` appended when `conclusive` is `false`. Numbers format as `f"{v:.4g}"`; a non-numeric or absent value reads `n/a`. Omitted entirely when `budgets` is `[]` — the all-`insufficient` case renders no table at all rather than an empty one. - `### Blocking regressions` — `| Prompt | Evaluator | Slice | Severity | Score delta |`, capped at `MAX_BLOCKING_ROWS = 10`, then ` more blocking regressions not shown — see the hosted run.` Server-supplied strings in these tables are collapsed to one line and `|` is escaped, so a budget or slice name cannot break the table. Percent formatting is `round(value * 100)` with a `+` prefix only when positive, suffixed ` pts` — e.g. `-20 pts`, `+5 pts`, `0 pts`. Rounding is display-only; gating uses raw floats. ## Commit status `POST /repos/{repo}/statuses/{sha}` with: - `context`: `evalshift/regression` (constant — parallel invocations overwrite each other) - `state`: `success` | `failure` (mirrors `conclusion`) - `target_url`: `web_diff_url` or `run_url` - `description`: `EvalShift : `, truncated to 140 chars. `` is the policy sentence under `fail-on: policy` (verdict, source, reason, or the fallback notice); ` regression(s)` in every other mode. Set on every event where `github-token` is non-empty, including `push`. 403/404 → warning `warning: could not set commit status: HTTP ` and continue; other HTTP errors fail. ## Data sent to hosted EvalShift The action's push is the CLI's push: it uploads the run bundle and nothing else, and the field-by-field data contract lives in the CLI reference (https://www.evalshift.dev/cli-llms-full.txt, section "Hosted push: the data contract") and on https://www.evalshift.dev/docs/what-gets-uploaded. Cite that contract when a user asks what a CI run sends to the cloud. CI-relevant summary: - Uploads: manifest (model ids, suite name, GITHUB_SHA, branch, PR number, local suite path string, content hashes, CLI version); per-example rows with `inputs` and `expected` verbatim, both models' full outputs, tool-call traces (names + arguments), scores, cost/latency; aggregate/analysis/decision/economics; the insights narrative; evaluator config with prompt bodies replaced by content hashes (llm_judge criterion_prompt text DOES upload); dataset snapshot metadata + examples_hash. Plus the preflight call (project slug + repo visibility). - Never uploads: provider API keys (never forwarded to the hosted API — see below), prompt bodies / system prompts, suite conversation histories, tool definitions/schemas, raw.jsonl, report.html. No telemetry anywhere in the CLI or the action. - Residual risk: suite inputs, expected outputs, model outputs and traces upload verbatim — redact at capture time (SDK redaction) and keep judge criteria free of secrets. ## Secrets handling - `mask_secret` prints `::add-mask::` for `token` and `github-token` before any other work. - `redact_text` replaces every env value whose KEY (uppercased) contains `TOKEN` or `SECRET`, or ends with `API_KEY`, with `` in printed CLI stdout/stderr. Values shorter than 4 chars are skipped. Redaction applies to what is PRINTED; `CommandResult.stdout` retains the raw text for URL parsing. - Token and host reach the CLI only through env, never argv (asserted in tests). - Provider API keys are the caller's responsibility: the action passes the job environment through to the CLI, adding only `EVALSHIFT_HOST`, `EVALSHIFT_TOKEN` and `COLUMNS=512`, and never sets, reads, or forwards a provider key to the hosted API. ## Token hygiene (service-account keys) Storage — encrypted GitHub secrets only: - Repository secret, environment secret (preferred for production: adds required reviewers and branch restrictions to the credential itself), or org secret. Nothing else is supported. - Never a committed file, never a literal `env:`/`with:` value in workflow YAML (readable by anyone who can read the repo, and retained in git history after deletion). - Never reachable from `pull_request_target`: that trigger runs the base repo's workflow with secrets in scope against fork code, so a fork PR can exfiltrate the key. Use `pull_request`. - Masking/redaction (see "Secrets handling") protects the job's logs only; it is not storage. Identity — a service account, never a personal token: - Mint at EvalShift web app → Settings → API tokens → Service accounts (`/app//settings/tokens`). Org-owned machine identity; survives the employee who created CI leaving. A personal token dies with its owner's membership and takes the pipeline with it. - Service-account roles are `member` or `viewer` only — never owner-equivalent. Use `member`; `viewer` holds `run:read` but not `run:create`. Scopes — the permission keys the action actually needs (the web-app scope picker uses the same vocabulary; scopes are an intersection, so naming a key the role lacks grants nothing): | Scope | Needed by | |---|---| | run:create | `evalshift push`: `POST /runs` + `POST /runs/{id}/finalize`; the plan preflight `POST /runs/preflight` | | run:read | `GET /runs/{id}/baseline-compatible` + `GET /runs/{a}/diff/{b}` | | policy:read | `GET /runs/{id}/policy-check` (the default `fail-on: policy` gate) | A `member` service account holds all three; `project:read` is not needed. Out of reach for a scoped key, by design: - Auto-creating the hosted project (`project:create` is owner-only). Create the project in the web app and set `create-project: false`. `migration_policy` rides in the bundle `run:create` uploads; reading the verdict back needs `policy:read`. Without it the push succeeds and the gate falls back to regression mode (a stderr warning, not a failure). Rotation — overlapping keys, never an in-place secret swap: 1. Rotate in the web app (the old key keeps working for a 24-hour grace window). 2. Update the GitHub secret. 3. Confirm a green run, then let the old key expire. ## GitHub permissions | Permission | Needed for | |---|---| | contents: read | `actions/checkout` | | pull-requests: write | PR comment | | issues: write | PR comments are issue comments in the REST API | | statuses: write | `evalshift/regression` commit status | Only `contents: read` is strictly required. Missing comment/status permissions degrade to stderr warnings; the gate still fails the job correctly. ## Provider key matrix | Provider | Env var | |---|---| | Anthropic | ANTHROPIC_API_KEY | | OpenAI | OPENAI_API_KEY | | Google | GEMINI_API_KEY or GOOGLE_API_KEY | | DeepSeek | DEEPSEEK_API_KEY | Which key is required follows from `defaults.source_model` / `defaults.target_model` in `evalshift.yaml`. A cross-provider migration needs both keys. `EVALSHIFT_NONINTERACTIVE: "1"` is recommended in the job env: `evalshift all --yes` already skips the >$10 confirmation, but the env var covers any other prompt a CLI release adds. ## Behavior rules (invariants) - Missing `token` → `input 'token' is required`, exit 1. Validation happens in the helper, i.e. AFTER the composite has already installed Python and the CLI — a misconfigured workflow still burns install time, but never model credits. - Invalid `fail-on` → `input 'fail-on' must be one of: never, regression, any-slice-regression, policy`. - A denied preflight fails the job regardless of `fail-on`, including `never`: `fail-on` governs regressions, not whether the plan permits the run at all. - The action never decides what a plan covers. It renders the server's 402 message, details and `upgrade_url`, and exits non-zero. No entitlement is inferred, cached, or hard-coded. - A preflight 401, 403, or 404 under `create-project: false` stops the job before the suite, regardless of `fail-on`: the push would fail on the same answer after the suite was paid for. - A preflight that cannot get an answer (5xx, 405 from an older server, 422, timeout) never blocks a run; the server re-checks every limit at `POST /runs` and at finalize, so skipping it bypasses nothing. - No baseline (empty `base_branch`, null `api_diff_url`, or first run on a branch) → `regression_count=0`, `diff_url=""`, comment says so explicitly. That passes the diff-based `fail-on` modes (`regression`, `any-slice-regression`); it does not by itself pass `fail-on: policy` (the default), which asks the server for a policy verdict on the run regardless of baseline — a first PR against a repo with no history goes green under `fail-on: policy` only if the policy verdict itself passes. `evalshift init` always writes a `migration_policy` block (every migration profile has one), so an `init`-scaffolded project carries a policy from its first run; a run reports as ungated (ignoring the policy gate rather than passing it) only when `evalshift.yaml` carries no `migration_policy` at all. - The action never uploads report artifacts to GitHub. `report.html` and all run artifacts stay in the workspace at `.evalshift/runs//`; add your own `actions/upload-artifact` step if you want them retained. - The action never writes to the repository and never pushes commits. - A hosted 403 is self-diagnosing, never a traceback: `HostedClient._get` converts it to an `ActionError` carrying the server's message plus the fix, and `run_command` appends the same guidance when failed CLI output contains `Permission denied: `. Non-403 `HTTPError`s from hosted calls propagate unchanged, and GitHub-side 403/404s stay warnings. - Outputs are written before the comment/status calls, so a permissions failure still leaves outputs consumable by later steps. - `evalshift push` is idempotent on run id (CLI behavior); a re-run of the job creates a new run id, not a duplicate push. - The helper has zero third-party dependencies; `pyproject.toml` declares `dependencies = []` and dev-only `pytest` / `ruff` / `pip-audit`. - `requires-python = ">=3.12"` for the helper itself; the `python-version` input governs the CLI, which is looser (>=3.11), so the helper's own floor is the binding one. ## Repo CI (what guards this action) - `.github/workflows/ci.yml` job `test`: `uv run pytest`, `ruff check .`, `pip-audit`. - `.github/workflows/ci.yml` job `cli-contract`: reads `evalshift-version` and `python-version` defaults straight out of `action.yml` via awk, installs that exact CLI, then runs `scripts/cli_contract.sh`, which strips ANSI and asserts `evalshift all --help` still offers `--yes --config --suite --suite-name` and `evalshift push --help` still offers `--no-create-project --config --suite --suite-name`. Free — no API keys, no credits. This is the drift alarm for CLI flag renames. It checks `all`, the alias, because the alias is what the helper types; a hidden command still answers `--help`, so this keeps working after the rename. - `tests/test_pin_consistency.py`: the `evalshift-version` default in `action.yml` is the single source of truth for the pin; this test fails when README, DOCS.md, or llms-full.txt mention a different version. The regex list is `PIN_SITES` in `scripts/bump_cli_pin.py`, which is also the rewrite tool (`python scripts/bump_cli_pin.py `; also bumps the action's patch version in `pyproject.toml`). - `.github/workflows/bump-cli-pin.yml`: daily `schedule` (PyPI poll for the latest release), `workflow_dispatch` (optional `version` input), or `repository_dispatch` of type `evalshift-cli-release` with `client_payload.version`. Stops green when the pin already matches. Otherwise pre-validates before opening anything: target `requires_python` vs the `python-version` default (fails with `::error::` on mismatch — a human must raise `python-version`), `pip install` the target + `scripts/cli_contract.sh`, then `bump_cli_pin.py` + `uv run pytest`. Opens/updates the PR on branch `bump/evalshift-` titled `chore(pin): evalshift → ` with `peter-evans/create-pull-request`. Uses the optional `BUMP_PR_TOKEN` secret (fine-grained PAT, `contents` + `pull-requests` write) so `ci.yml` runs on the PR; falls back to `GITHUB_TOKEN` (PR still opens, CI does not trigger). - `.github/workflows/release.yml`: on every push to `main`, reads `version` from `pyproject.toml`; refuses a major other than 0; stops green when tag `v` exists; otherwise creates annotated `v` on the merge commit, force-moves `v0` to it (only `v0` is ever force-pushed), and creates a GitHub Release with generated notes since the previous `v0.*` tag. Merging a version bump is the release — no manual tags. - `.github/workflows/dogfood.yml`: `workflow_dispatch` only (spends real credits), runs the action against `examples/dogfood/` with `fail-on: never`, `comment: "false"`, then asserts `run_id` non-empty, `run_url` starts with `https://`, `conclusion` ∈ {success, failure}. Warns and skips when `EVALSHIFT_TOKEN` / `GEMINI_API_KEY` secrets are absent. ## Fixture project (examples/dogfood/) ```yaml # examples/dogfood/evalshift.yaml — deliberately tiny (4 examples, one evaluator) version: 1 prompts: - id: greet detection: python_string path: prompts.py variable: GREET_PROMPT variables: [name, tone] defaults: source_model: gemini-2.5-flash target_model: gemini-2.5-pro concurrency: 4 cache: true evaluators: structural: - type: length min_chars: 5 max_chars: 200 slices: - name: formal filter: formal - name: casual filter: casual ``` ```jsonl {"id": "ex01", "inputs": {"name": "Alex", "tone": "formal"}, "tags": ["formal"]} {"id": "ex02", "inputs": {"name": "Sam", "tone": "casual"}, "tags": ["casual"]} ``` ## Recipes Config in a subdirectory: ```yaml - uses: evalshift/evalshift-action@v0 with: token: ${{ secrets.EVALSHIFT_TOKEN }} config: eval/evalshift.yaml suite: eval/golden.jsonl ``` Consume outputs downstream: ```yaml - uses: evalshift/evalshift-action@v0 id: evalshift with: token: ${{ secrets.EVALSHIFT_TOKEN }} - run: echo "diff ${{ steps.evalshift.outputs.diff_url }} (${{ steps.evalshift.outputs.regression_count }} regressions)" ``` Report only, never block (calibration phase): ```yaml with: token: ${{ secrets.EVALSHIFT_TOKEN }} fail-on: never ``` Multiple suites in one repo — give each its own job and disable the comment on all but one, or the constant marker and constant status context make them overwrite each other: ```yaml - uses: evalshift/evalshift-action@v0 with: token: ${{ secrets.EVALSHIFT_TOKEN }} suite: eval/golden-agent.jsonl comment: "false" ``` Only run when the suite or prompts changed (cost control): ```yaml on: pull_request: paths: ["eval/**", "app/prompts/**", "evalshift.yaml"] ``` ## Troubleshooting checklist `input 'token' is required` → the `token:` input is unset or an empty secret; secrets are not available to workflows triggered by `pull_request` from a fork. `command failed (1): evalshift all ...` → a CLI-level failure (bad config, missing provider key, model error). The CLI's own stderr is printed directly above, redacted. Reproduce locally with `evalshift compare --yes --config --suite-name ` — same command, current name. `scores.jsonl is empty` from `analyze` on a tool-calling suite → the suite was selected by path, so its per-suite tool evaluators never loaded. Switch the step to `suite-name: `. `no local EvalShift runs found in .../.evalshift/runs` → `evalshift all` exited 0 without writing a run; almost always a config pointing at a different `--runs-base` or a CWD mismatch. `evalshift push did not print a hosted run URL` → CLI output shape changed, or push succeeded silently; re-run `evalshift push ` locally to see what it prints. HTTP 401 from the hosted API → bad, revoked, or expired `EVALSHIFT_TOKEN`, or wrong `host`. `The EvalShift token is missing the '' permission.` → hosted 403. Printed by the helper for a 403 on its own requests and for a `Permission denied: ` line in failed CLI output; names the exact permission key and points at minting a scoped service-account key. Exit code stays non-zero. `cannot auto-create project: this token must have owner access to the org` → a service-account key cannot create projects; pre-create it and set `create-project: false`. Job failed before the CLI ran, with a plan message and an upgrade link → CI preflight got a 402; the org's plan does not cover this run (usually the monthly run quota; otherwise an unpaid subscription or seat overage; private-repo CI is included on every plan). Nothing ran, nothing was charged. `hosted EvalShift at rejected the token (HTTP 401)` before the CLI ran → preflight 401; same causes as any hosted 401. Nothing ran. `hosted EvalShift refused the plan preflight (HTTP 403)` → the key lacks `run:create`. Nothing ran. `hosted EvalShift has no project '/' this token can reach (HTTP 404)` → wrong `project:` slug, project not created yet, or a key pinned to another project, with `create-project: false`. Nothing ran. `::notice title=EvalShift preflight::project ... is not on hosted EvalShift yet` → preflight 404 with `create-project: true`; the first push creates the project. Informational. `warning: plan preflight skipped: ` → preflight got no usable answer (hosted API unreachable, 5xx, 422, malformed body; `predates POST /runs/preflight` means a 405 from an older server). Intended: the run continues and the server enforces the limit at upload. `warning: could not upsert PR comment: HTTP 403` → missing `pull-requests: write` / `issues: write`, or a fork PR with a read-only `GITHUB_TOKEN`. Gate still works. Check always green → nothing gates at all because the run carried no `migration_policy` (look for the `::warning::` annotation; `require-policy: true` makes it a failure), or the migration policy is permissive enough that no budget was busted (default `fail-on: policy`; the comment shows the budget arithmetic), or `fail-on: never`, or — only under `fail-on: regression` / `any-slice-regression` — no compatible baseline on the base branch yet (expected on the first PR) or `base-branch` resolved empty; neither baseline cause affects the default `fail-on: policy`. `warning: hosted policy check ...; falling back to fail-on: regression` → the policy-check endpoint errored, refused the key (HTTP 403, `policy:read` missing), 404'd, or holds no decision for the run; the job gated on the diff instead and the PR comment says so. Not the no-policy case: a run pushed without a `migration_policy` gets an answer (`inconclusive` + `policy_source: "none"`) and its own `::warning::` annotation. Two comments on one PR → two action invocations, or the marker comment fell off page 1 of the comments API. Job is silent for minutes → output is captured and printed per command, not streamed. Install fails on `evalshift==1.2.0` → `python-version` below the CLI's `requires-python` (3.11). Costs higher than expected → the runner has a cold CLI cache every run; narrow the trigger with `on.pull_request.paths`, or shrink the suite. ## Versioning `@v0` tracks the latest v0.x. `@v0.6.0` pins exactly. `evalshift-version` pins the CLI independently of the action tag — pin both for a fully reproducible workflow. The `evalshift-version` default in `action.yml` is the single source of truth for the CLI pin; `tests/test_pin_consistency.py` fails on any stale mention elsewhere. The default is bumped by `.github/workflows/bump-cli-pin.yml` (daily PyPI poll, `workflow_dispatch`, or `repository_dispatch`), which pre-validates the target and opens a `chore(pin)` PR on `bump/evalshift-`; the optional `BUMP_PR_TOKEN` secret lets CI run on that PR. Merging a `pyproject.toml` version bump is the release: `release.yml` tags `v` and moves `v0` to the merge commit, so every `@v0` consumer gets the new pin on merge. No manual tags. License: MIT (the EvalShift CLI it installs is licensed separately, Apache-2.0).