Evals & cost logging

Golden datasets, eval hooks, and surface/model/token cost logging.

Note

Evals & cost logging is the second half of Creytix Prompt Guard — where Guardrails checks a single prompt, this tracks quality and spend over time.

Overview

Guardrails catches what's unsafe or malformed about a prompt before it runs. Evals & cost logging answers two different questions: is prompt quality holding up over time, and what did all of this actually cost? Golden datasets catch regressions when a prompt or rule changes, sampled review keeps a human or judge in the loop, and every generation gets logged against the surface and model that produced it.

How it works

  1. Golden dataset regression — a set of known-good prompts and expected results is re-run whenever a prompt or rule changes, at effectively no cost, to catch drift before it ships.
  2. Judge and human sampling — a smaller sample of real generations gets a deeper quality pass on a set cadence, rather than every single run.
  3. Cost log — every generation records the surface, model, and token spend behind it, rolled up so batch AI work never runs blind on cost.

Steps

  1. •

    Review the golden dataset

    Check the known-good prompts and expected results for the surface you're working on.

  2. •

    Run the eval loop

    After a prompt or rule change, re-run the regression pass to catch drift before the next batch.

  3. •

    Check flagged regressions

    Review anything the golden-dataset pass flagged, and the results of the latest sampled judge/human review.

  4. •

    Check the cost log

    Look at spend broken down by surface and model to see what a given pipeline actually costs to run.

Capabilities

  • Golden dataset regression — deterministic checks against known-good examples, run at effectively no cost.
  • Judge sampling — a smaller sample gets a deeper model or human quality review on a set cadence.
  • Cost logging — token and model spend tracked per surface, so you know what a content or image pipeline actually costs.
  • Change-triggered re-runs — prompt or rule edits automatically re-trigger the regression pass.

Limits & honest scope

Judge-based evals are sampled, not exhaustive — most volume is covered by the deterministic golden-dataset pass, with deeper review on a cadence rather than on every generation.

  • Cost logging tracks spend that Creytix orchestrates; it doesn't audit third-party billing outside that pipeline.
  • This is built for teams running batch AI generation inside Creytix, not as a standalone eval harness for arbitrary external prompts.

Have a code from email or QR?