Concepts
Quality checks
Why you need output validation regardless of muster — and how muster amplifies it.
Documentation Index
Fetch the complete documentation index at: /llms.txt
Use this file to discover all available pages before exploring further.
Why you need output validation regardless of muster — and how muster amplifies it.
# This code belongs in your codebase regardless of muster
line_items_sum = sum(item["total"] for item in result["line_items"])
subtotal_ok = abs(line_items_sum - result["stated_subtotal"]) < 0.01
# Same check. One extra line to tell muster about it.
muster_emit(job_id, checks=[{
"check_id": "subtotal_arithmetic",
"severity": "HIGH",
"passed": subtotal_ok,
}])
result = llm.extract(invoice_pdf_text)
checks = [
# Did we get anything?
{"check_id": "output_not_empty", "severity": "HIGH", "passed": bool(result)},
# Do the line items add up?
{"check_id": "subtotal_arithmetic", "severity": "HIGH", "passed": line_sum_ok,
"expected": str(stated_subtotal), "actual": str(computed_sum)},
# Does subtotal + tax = grand total?
{"check_id": "grand_total_arithmetic", "severity": "HIGH", "passed": grand_total_ok},
# Is the date format valid?
{"check_id": "invoice_date_valid", "severity": "MEDIUM", "passed": date_ok},
]
decision = agent.decide(application)
checks = [
# Is the decision one of the expected values?
{"check_id": "decision_is_valid_enum", "severity": "HIGH",
"passed": decision in {"APPROVE", "REJECT", "ESCALATE"}},
# Did it explain itself?
{"check_id": "decision_has_rationale", "severity": "HIGH",
"passed": len(decision_rationale) > 50},
# Did it not refuse to answer?
{"check_id": "no_refusal_in_output", "severity": "HIGH",
"passed": "cannot" not in decision.lower() and "unable" not in decision.lower()},
]
summary = agent.summarise(document)
checks = [
{"check_id": "output_not_empty", "severity": "HIGH", "passed": bool(summary)},
# Is the summary meaningfully shorter than the source?
{"check_id": "compression_achieved", "severity": "MEDIUM",
"passed": len(summary) < len(document) * 0.4},
# Does it avoid hallucinated citations?
{"check_id": "no_fabricated_urls", "severity": "HIGH",
"passed": not contains_urls_not_in_source(summary, document)},
]
| Without muster | With muster |
|---|---|
| Log files per agent | Fleet-wide heatmap — all agents, all checks, one view |
| Manual review to spot degradation | Automatic alerts when pass rate drops by >15% |
| No idea which check is failing most | Sorted by worst-performing check across all agents |
| No external reference point | Benchmark comparisons against similar agents (opt-in) |
| Ops team reads logs reactively | Finance team catches invoice errors before payment runs |