llms.txt Content
# Metrx
> Metrx heads every AI workload toward its best proven configuration. It finds cheaper and better configurations, proves them against a randomized holdout, and switches on your say-so — as models and prices change.
## What Metrx Does
- Generates candidate configurations per workload: model, prompt, parameters, routing
- Proves them with pre-registered randomized trials on your traffic, once it has evidence
- Judges every candidate against an acceptance contract you approve, not a generic score
- Can put winners live under the authority you define, per change class
- Watches live configurations for drift and re-verifies incumbents on model releases
- Rolls back on contract violation and records the event
- Keeps a full evidence ledger — wins, losses, and rollbacks alike
- Tracks per-agent, per-model LLM cost and token use across providers
## What Metrx Does NOT Do
- Metrx does NOT store prompt or completion content — only metadata, cost signals, and the outcome signals you choose to send
- Metrx is NOT a prompt management tool — use a prompt registry for versioning
- Metrx is NOT an agent hosting platform — it works with your existing stack
- Metrx does NOT claim a population benchmark is your result — every number carries its provenance
## Procedural Integrity
Metrx optimizes and also measures, so the trust claim is about procedure, not about neutrality:
1. Candidate generation is separated from evaluation
2. You approve the objectives and the acceptance contract; optimization happens inside them
3. Treatment is randomized, so measured effects are causal rather than before/after coincidence
4. Negative and inconclusive results are preserved — never retried until positive
5. Every promotion carries a reproducible evidence record with provenance and exclusions
## Provenance Labels
Every number on every Metrx surface carries one of these:
- Measured — directly observed from production data
- Attributed — a modeled estimate with stat