open source · reliability stack
axon labs
tools for when an llm has to be right enough to ship. detection, observability, context, chunking — no theater.
hallucination detection. four methods — semantic entropy, log probs, web grounding, llm-as-judge. one api call. a confidence score you can actually ship on.
llm observability. decorator-based tracing, behavior evaluation, security scanning, multi-channel alerts. deep datadog integration.
intelligent context builder. pattern-based search with token-aware windowing. 85% token reduction at 99.2% quality retention in evals.
semantic double-pass chunking for rag. boundaries by meaning, not character count. lookahead merging keeps related content together.
pii detection and redaction via gliner. 150+ entity types. reversible — store mappings, restore originals from redacted text.
wearable ai second brain. raspberry pi zero w + whisper + mistral. records conversations, indexes them, surfaces them when needed.